CS408 — Midterm Summary (Lectures 1–22)
📘 Lecture 1 — Introduction to Human Computer Interaction – Part I
📖 Overview: This lecture introduces the field of Human Computer Interaction (HCI) by examining how technology has evolved from the 1950s to the present day. It uses compelling real-world examples to demonstrate why HCI is critically important for designing systems that serve human needs rather than frustrate them.
🗂️ Topics Covered
The lecture covers the historical evolution of computers from expensive, specialist-only machines to ubiquitous everyday devices, the consequences of poor interface design illustrated through multiple case studies (airline crash, digital cameras, alarm clocks, cars, and warships), and concludes with a formal definition of HCI and its role in bridging the gap between human understanding and machine interfaces.
📝 Lecture Summary
Introduction to Human Computer Interaction – Part I
The lecture opens by describing how computers have invaded every aspect of modern life. Devices like ATMs, cell phones, VCRs, remote controls, ticketing machines, digital personal organizers, calculators, watches, photocopiers, toasters, banks, air conditioners, microwaves, and medical equipment all contain computers. Unlike the 1950s when only highly skilled technical specialists used computers, today users come from every walk of life — commerce, farming, education, retailing, defense, manufacturing, and entertainment.
In the 1950s, computers were extremely difficult to use for several reasons: they were very large and expensive machines (making human labor cheap by comparison), used only by technical specialists familiar with punch cards, and little was known about how to make them easier to use. None of these conditions hold today — computers have become much less expensive, users are diverse, and we understand much more about fitting machines to people's needs.
The silicon chip enabled dramatic decreases in computing costs, allowing miniaturization and packing large numbers of circuits onto tiny chips. In less than thirty years, computers changed from huge machines in air-conditioned rooms to small portable devices. The development of the first personal computers in the 1970s was a major landmark, providing interactive computing power for individual users at low cost.
💡 Why this matters: Computers now perform more and more tasks, making life without them as unimaginable as life without electricity. However, this invasion has both bright and dark sides.
Bright side results: Computers are enabling new discoveries, leading to efficiencies, and making our life easy and convenient.
Not so bright side results: Computers are annoying us, infuriating us, and even killing a few of us. We are irreversibly dependent on these "hopeful monsters," so we must fundamentally rethink how human and machines interact.
1.1 Riddles for the Information Age
This section presents four case studies demonstrating what happens when computers are integrated into everyday objects.
Riddle 1: Computer with an Airplane In December 1995, American Airlines Flight 965 departed from Miami to Cali, Columbia. On landing approach, the pilot of the 757 needed to select the navigation fix "ROZO". He entered "R" into his navigation computer, which returned a list of nearby fixes starting with "R". The pilot selected the first option — "ROMEO", 132 miles northeast instead of "ROZO". The jet descended into a north-south valley, followed the flight computer indications, began an easterly turn, and slammed into a granite peak at 10,000 feet. 152 passengers and all 8 crewmembers perished; only 4 survived with serious injuries.
Riddle 2: Computer with a Camera The evolution of cameras demonstrates increasing complexity:
- Thirty years ago: A 35mm Pentax Model H had a small battery for the light meter, simple to maintain
- Fifteen years ago: A 35mm Canon T70 used two AA batteries with a simple On/Off switch
- Five years ago: A first-generation digital camera had automatic shutdown after one minute of inactivity
- One year ago: A Panasonic PalmCam had an Off/Rec/Play switch with modes
- Newest: A Nikon CoolPix 900 (third-generation) displays a Windows-like hourglass while booting up, with four switch settings: Off/ARec/MRec/Play
🔑 Definition — ARec: "automatic record" mode on the digital camera 🔑 Definition — MRec: "manual record" mode on the digital camera
The Nikon's power management program causes problems: when zooming, charging flash, and energizing display simultaneously drains power, it suspends picture-taking ability. When the user presses the shutter button, it shuts down the LCD display to shed load. When the user drops the camera, it senses more power available and takes a picture — of the user's kneecap.
Riddle 3: Computer with an Alarm Clock A JVC FS-2000 clock-radio has sophisticated computer features including fading up volume at 6 AM. However, it's very hard to tell when the alarm is armed, so it occasionally fails to wake the user on Monday and wakes them early on Saturday. The alarm indicator is a small symbol in the upper left corner of an alphanumeric liquid crystal display (LCD) — barely visible in a dim bedroom. The backlight only comes on when CD or radio is explicitly turned on, but the alarm won't sound while the CD is left on regardless of alarm settings.
To disarm the alarm: press "Alarm" button once. To arm it: press exactly five times (showing alarm time, sound-off time, radio/CD selection, preset volume, then normal view with alarm armed). One additional press disarms it.
By contrast, an old non-computerized alarm clock had a single red light when armed, dark when not armed — simple and clear.
Riddle 4: Computer with a Car Porsche's Boxster has seven computers, including one dedicated to engine management. In early models, if fuel level got very low (about one gallon remaining), centrifugal force from a sharp turn could cause fuel to collect at the tank side, letting air enter fuel lines. The computer interpreted this as catastrophic injection system failure and shut down the ignition completely, preventing restart until the car was towed and serviced. The only solution Porsche offered was disconnecting the battery for five minutes to reset the computer.
Riddle 5: Computer with a Warship In September 1997, the USS Yorktown (an Aegis guided-missile cruiser) stopped dead in the water during fleet maneuvers. A Navy technician entered a zero while calibrating an on-board fuel valve into a Pentium Pro running Windows NT. The program attempted to divide another number by zero (an undefined mathematical operation), causing a complete crash of the entire shipboard control system. The engine halted, and the ship sat for 2 hours and 55 minutes until towed to port.
1.2 Role of HCI
Humans designed the frustrating interfaces we hate. Humans continue to use dysfunctional machines even as awkward interfaces strain their eyes, ache their backs, and ruin their wrist tendons. HCI plays a role to bridge the gap between machine interfaces and human understanding.
🔑 Definition — Human-Computer Interaction (HCI): "A discipline concerned with the design, evaluation and implementation of interactive computing systems for human use and with the study of major phenomena surrounding them" — ACM/IEEE
⭐ Key Takeaways
HCI is fundamentally about designing systems that fit human needs rather than forcing humans to adapt to machine logic. The lecture demonstrates through real-world tragedies (plane crash killing 160 people) and everyday frustrations (cameras, alarm clocks, cars) that poor interface design has serious consequences — from inconvenience to loss of life. The definition of HCI encompasses design, evaluation, implementation, and study of interactive systems. Computers will continue to infiltrate every product because it's economically cheaper than mechanical methods, making HCI knowledge essential. The fault for bad design lies with humans, not machines, and we must fundamentally rethink human-machine relationships.
🧠 Quick Revision Questions
- What four factors made 1950s computers extremely difficult to use?
- What specific interface failure caused the crash of American Airlines Flight 965?
- How does the Nikon CoolPix 900's power management program fail its user?
- What happened to the USS Yorktown and why?
- According to the ACM/IEEE definition, what four components comprise the discipline of Human-Computer Interaction?
📘 Lecture 2 — Introduction to Human-Computer Interaction – Part II
📖 Overview: This lecture examines the adverse impacts of poor human-computer interaction through real-world examples like airplane crashes and confusing car interfaces. It introduces the formal definition of HCI, contrasts human and computer capabilities, and explains the concept of "software apartheid" — how poorly designed software creates social and economic divisions. The lecture also differentiates the user-centered focus of HCI from the system-centered focus of Software Engineering.
🗂️ Topics Covered
The lecture covers the formal ACM/IEEE definition of HCI, analyzes the fatal American Airlines Flight 965 crash as a case study of poor computer communication, and examines other problematic interfaces like BMW's iDrive system and ATMs. It then contrasts the nature of humans versus computers, introduces the concept of software apartheid, and concludes by differentiating the approaches of Software Engineering (system-centered) and HCI (user-centered).
📝 Lecture Summary
2.1 Definition of HCI
The formal definition of Human-Computer Interaction is provided by the ACM/IEEE: “Human-Computer Interaction is a discipline concerned with the design, evaluation and implementation of interactive computing systems for human use and with the study of major phenomena surrounding them.” This establishes HCI as a discipline that goes beyond just building systems to understanding the entire context of human use.
🔑 Definition — HCI (Human-Computer Interaction): A discipline concerned with the design, evaluation, and implementation of interactive computing systems for human use and with the study of major phenomena surrounding them.
2.2 Reasons of non-bright Aspects
This section examines real-world examples where computer systems failed their human users, often with tragic consequences. The core argument is that computers "may tell us facts but they don't inform us."
Airplane + Computer Case Study: The lecture analyzes the fatal crash of American Airlines Flight 965 in Cali, Colombia. The National Transportation Safety Board (NTSB) declared the cause "human error" — the pilot selected the wrong navigation fix (ROMEO instead of the correct fix for landing). However, the lecture argues this was not truly the pilot's fault. The flight computer's front panel showed the selected fix and a course deviation indicator (needle), but the needle looked the same whether the plane was on course to land or on course to crash. The computer gave no indication that ROMEO was an inappropriate fix for the landing approach.
💡 Why this matters: The computer was "utterly unconcerned with the actual flight and its passengers" — it only cared about its own internal computations, not the real-world consequences of its output.
The Computer Industry Joke: A pilot lost in clouds asks "Where am I?" and gets the reply: "You are in an airplane about 100 feet above the ground." The pilot immediately knows the course because the answer was "completely correct and factual, yet it was no help whatsoever" — revealing the responder was a Microsoft software engineer. This joke highlights a "fundamental truth about computers: They may tell us facts but they don't inform us. They may guide us with precision but they don't guide us where we want to go."
I-Drive Car Device: BMW's iDrive system uses a single large, multifunction knob in the console between the front seats. Users rotate the knob to scroll through menus, push it axially to select, and move it forward/back or side-to-side to change functions (communications, climate control, navigation, entertainment). Several hundred functions are controlled with this one device.
The fundamental problems include:
- Users must look at an LCD screen to see what they're selecting, taking their eyes off the road
- Force feedback indicates the menu level but not the actual content in the menus
- It "takes 15 minutes to change a Radio Channel"
- Results in constant calls to help desk
Feature Shock: "Every digital device has more features than its manual counterpart, but manual devices are easier to use." Hi-tech companies add more features thinking it improves the product, but the product becomes complicated. The key insight: "Bad process can't improve product."
Computer + Bank (ATM Case Study): The ATM exhibits "sullen and difficult behavior." If the user makes the slightest mistake (like selecting "savings" when they only have checking), it rejects the entire transaction and forces them to start over. The ATM knows the user doesn't have a savings account, yet still offers it as a choice — then punishes the user for selecting it. "The only difference between me selecting 'saving' and the pilot of Flight 965 selecting 'ROMEO' is the magnitude of the penalty."
The lecture notes that this behavior "is not intrinsic to computers... nothing is intrinsic to computers: they merely act on behalf of their software, the program." Programs are "as malleable as human speech" — they can be polite or rude. "All it takes is for someone to describe how. Unfortunately, programmers aren’t very good at teaching that to computers."
💡 Why this matters: This problematic behavior across all these examples is what drives the need for the emerging field of Human-Computer Interaction.
2.3 Human verses Computer
Human species: Human beings are the "most interesting and fascinating specie on planet." They are complex, diverse, intelligent, and free in nature. Humans "think on a problem dynamically" and can find solutions that "may not exist before." They can invent. They are both rational and emotional. "And fortunately or unfortunately they make mistakes."
Computer species: Computers are an invention of human beings. They are "complex but also pretty dumb." They "can think but can't think on its own will" — they think only as directed. Their speed is "marvelous." They "do not tire. It is emotionless. It has no feelings, no desires." And "they do not make mistakes."
Before vs. Now: Before computer penetration in daily life, humans performed tasks with personal responsibility and interaction. A salesperson could judge a client's mood through tone, attitude, and body language, providing relevant answers. Now, "in this age of information technology we are expecting computers to mimic human behavior" — e.g., eCommerce websites acting as salespersons. "That is now; a dumb, unintelligent and inanimate object will perform the complex task which was performed by some human being."
2.4 Software Apartheid
Apartheid definition: "Racial segregation; specifically: a policy of segregation and political and economic discrimination against non-European groups in the Republic of South Africa."
🔑 Definition — Software Apartheid: Institutionalizing obnoxious behavior and obscure interactions of software-based products.
The lecture argues that programmers "work in high-tech environments, surrounded by their technical peers" (like Silicon Valley), with limited contact with frustrated users. The term "computer literacy" is used as a requirement — assuming users must acquire fundamental training to use computers. However, "having a computer literate customer base makes the development process much easier... but it hampers the growth and success of the industry and of society."
The Car Analogy Refuted: Apologists counter that "you must have training and a license to drive a car," but they overlook that "a mistake with software generally does not [cause death]. If cars were not so deadly, people would train themselves to derive the same way they learn excel."
Social Division: The requirement for computer literacy creates "a demarcation line between the haves and have-nots in society." If you must master a computer to succeed in the job market, the difficulty of using interactive systems "forces many people into menial jobs rather than allowing them to matriculate into more productive, respected and better-paying jobs."
Key Arguments:
- "Users should not have to acquire computer literacy to use computer for common, rudimentary task in everyday life"
- An accountant "trained in the general principles of accounting, should not have to become computer literate to use a computer in her accounting practice. Her domain knowledge should be enough to see her through"
- The term "computer literacy" becomes "a euphemism for social and economic apartheid"
- "Software-based products not INHERENTLY hard to use — Wrong process is used to develop them"
🔑 Definition — Computer Literacy (as critiqued): A key phrase that brutally bifurcates our society, creating an artificial barrier between those who can use technology and those who cannot, based on technical knowledge that should be inconsequential.
Software Engineering and HCI
"There is a basic fundamental difference between the approaches taken by software engineers and human-computer interaction specialists. Human-computer interface specialists are user-centered and software engineers are system-centered."
Software Engineering Approach:
- Good at modeling certain aspects of the problem domain
- Formal methods developed to represent data, architectural, and procedural aspects
- Deals with managerial and financial issues well
- Useful for specifying and building functional aspects of a software system
HCI Approach:
- Emphasizes developing a deep understanding of user characteristics
- Clear awareness of the tasks a user must perform
- Tests design ideas on real users
- Uses formal evaluation techniques to replace intuition in guiding design
- "This constant reality check improves the final product"
⭐ Key Takeaways
The most critical lesson is that computers can provide precise, factually correct information while being tragically unhelpful or dangerous — they may tell us facts but don't inform us. Poor interface design has real-world consequences ranging from fatal accidents (Flight 965) to everyday frustration (ATMs, car controls). The concept of software apartheid reveals that requiring "computer literacy" for basic tasks creates artificial social and economic divisions, forcing people to master irrelevant technical distinctions (like RAM vs. Hard Disk) to participate in the modern economy. The fundamental difference between SE and HCI is that SE is system-centered (focused on data, architecture, and procedures) while HCI is user-centered (focused on understanding user characteristics, tasks, and testing with real users). Poor software interfaces are not inevitable — they result from using the wrong development process.
🧠 Quick Revision Questions
- What is the ACM/IEEE definition of Human-Computer Interaction?
- Why does the lecture argue that the crash of Flight 965 was not truly "human error" by the pilot?
- What is "software apartheid" and how does the concept of "computer literacy" contribute to it?
- Compare and contrast the nature of human beings versus computers as described in this lecture.
- How does the approach of Software Engineering differ from the approach of HCI specialists in developing interactive systems?
📘 Lecture 3 — Introduction to Human-Computer Interaction – Part III
📖 Overview: This lecture explores the real-world consequences of poor design in high-tech tools, using vivid examples from the airline industry and studies on "techno-rage." It explains why the high-tech industry is in denial about usability problems, and defines what it takes to succeed in the new economy by mastering the intersection of business and technology.
🗂️ Topics Covered
The lecture begins by examining the negative effects of bad tools, using case studies of in-flight entertainment systems. It then discusses the high-tech industry's denial of usability failures and the phenomenon of "techno-rage," supported by survey data on user frustration and attacks on computers. Finally, it presents success criteria for the new economy, emphasizing the need for business-savvy technologists and technology-savvy businesspeople, and concludes with a legal case about web accessibility.
📝 Lecture Summary
Effect of Bad Tools
The lecture uses the example of in-flight entertainment (IFE) systems to illustrate the catastrophic consequences of bad design. One airline’s IFE was so frustrating for flight attendants that they preferred shorter, less glamorous routes to avoid using it. This caused a serious morale problem, costing the airline money, customer loyalty, and staff loyalty.
Another airline’s IFE was even worse because it linked movie delivery with cash collection. This forced flight attendants to walk to a console, enter a password, and complete a cash-register transaction before a passenger could watch a movie. Out of sheer frustration, flight attendants would trip the circuit breaker on the system at the start of each flight and announce it was broken. The airline had spent millions on a system so obnoxious that its users deliberately turned it off to avoid interacting with it, causing a catastrophic financial loss.
💡 Why this matters: The software inside the IFEs worked with flawless precision but was a resounding failure because it misbehaved with its human keepers. This highlights that technical perfection is worthless without good user experience.
3.1 An Industry in Denial
The high-tech industry is in denial about the fact that its products are difficult to use. Technologists and software engineers believe they have done their best and that only new technology (like voice recognition or AI) can improve the user experience. However, the real problem is not technology, but a deficiency in culture, training, and attitude in the development process.
The industry has inadvertently put programmers and engineers in charge, whose hard-to-use engineering culture dominates. The lecture states, "We have let the inmates run the asylum." When creators examine their products, they see rich features, but ignore how difficult they are to use, how long they take to learn, and how they degrade the people who must use them.
3.2 Techno-Rage
A widespread undercurrent of techno-rage is rising. An article in the Wall Street Journal described a popular video clip of a man violently attacking his computer, which tapped into a powerful feeling of frustration. A joke about "Computer Tourette’s" describes otherwise-normal people swearing repeatedly at their monitors due to frustrating interactions.
The Novatech survey of 4,200 replies revealed a dark story: one in every four computers has been physically attacked by its owner. Technical support people worldwide have stories of computers brought in for repair that were deliberately smacked, kicked, or otherwise injured.
Another study by Symantec and Britain’s National Opinion Poll found that over 40% of British users have sworn at, kicked, or abused their computers. A separate survey by Concord Communications found that 83% of U.S. respondents had witnessed such attacks. Stress from computer rage has resulted in a loss of productivity and is described as being "much more prolific than road rage." The most frustrating incident is when people lose their work after a computer crash.
Key statistics from the studies:
- 70% swear at PCs
- 67% experienced frustration, exasperation, and anger
- Nearly half of all computer users had become angry at some time.
3.3 Success Criteria in the New Economy
The successful professional of the 21st century is either a business-savvy technologist or a technology-savvy businessperson.
- A technology-savvy businessperson knows their success depends on the quality of information and how they use it.
- A business-savvy technologist is an engineer with a keen business sense.
All businesspeople can be divided into two categories: those who master high technology, and those who will soon go out of business. Business is now information processing, and digital information is the beating heart of the workday. Bad information processes can cause a company to lose everything. The lecture shares several facts and figures highlighting the cost of bad user experience:
- Users can only find information 42% of the time (Jared Spool).
- 62% of web shoppers give up looking for an item (Zona Research).
- 50% of potential sales from a site are lost because people can't find items (Forrester Research).
- 80% of software lifecycle costs occur after release, and 80% of that is due to unmet user requirements (IEEE Software).
- 63% of software projects exceed cost estimates due to user changes, overlooked tasks, and poor communication (Communications of the ACM).
- BOO.com, a $204m startup, failed (BBC News).
- Poor web sites will kill 80% of Fortune 500 companies within a decade (Jakob Nielsen).
The product with bad user experience deserves to die. The lecture illustrates this with a table comparing two scenarios of an e-commerce site with a $100 million revenue potential. In Scenario A (good user experience), 0% sales are lost, and actual revenue is $100 million. In Scenario B (bad user experience), 50% of sales are lost, and actual revenue is only $50 million, resulting in a $50 million loss.
3.4 Computer + Information
The lecture asks, "What do you get when you cross a computer with information?" and answers by describing a lawsuit. Before the 2000 Sydney Olympics, a blind man, Bruce Lindsay Maguire, filed a lawsuit against the Sydney Organizing Committee for the Olympic Games (SOCOG) for unlawful discrimination. He complained that the website was not accessible to him.
According to the law in many countries, organizations must ensure their websites are accessible to disabled persons. The SOCOG website was not. The result was that the complainant won the case and was awarded damages. This was very embarrassing for SOCOG and the company that developed the website.
⭐ Key Takeaways
Bad tools can cause catastrophic financial and morale damage, as shown by the IFE examples where users deliberately turned off expensive systems. The high-tech industry is in denial, blaming a lack of new technology instead of a flawed development culture that prioritizes features over user-friendliness. This has led to widespread "techno-rage," with a significant percentage of users verbally and physically attacking their computers, costing productivity and increasing stress. To succeed in the new economy, professionals must be both business-savvy and tech-savvy, as poor user experience leads to massive revenue losses, project failures, and even legal consequences for accessibility violations.
🧠 Quick Revision Questions
- What were the two specific design flaws in the in-flight entertainment systems that led flight attendants to sabotage them?
- According to the lecture, what is the true cause of the difficulty in using high-tech products, and who is blamed instead?
- What was the result of the Novatech survey regarding physical attacks on computers, and what did the Symantec study find about the prevalence of "techno-rage"?
- Name three statistics that prove the significant financial cost of a bad user experience on e-commerce websites.
- What was the legal outcome of the Bruce Lindsay Maguire vs. SOCOG case, and what core accessibility issue did it address?
📘 Lecture 4 — Goals & Evolution of Human Computer Interaction
📖 Overview: This lecture defines the core goals of Human-Computer Interaction, distinguishing between usability goals (effectiveness, efficiency, safety, utility, learnability, memorability) and user experience goals (satisfaction, enjoyment, fun). It also traces the historical evolution of HCI through landmark systems like the Dynabook, Xerox Star, and Apple Lisa, highlighting key design principles that shaped modern interfaces.
🗂️ Topics Covered
The lecture begins with a formal definition of HCI from ACM/IEEE, then delves into the two major categories of HCI goals: usability goals and user experience goals. Each usability goal is explained in detail with real-world examples. The second half covers the evolution of HCI, focusing on three pioneering systems—Dynabook, Star, and Lisa—and their contributions such as direct manipulation, WYSIWYG, and consistency. The section concludes with transatlantic research differences and the emergence of the user interface concept.
📝 Lecture Summary
4.1 Goals of HCI
The term Human Computer Interaction (HCI) was adopted in the mid-1980s to describe a field broader than just interface design, encompassing all aspects of interaction between users and computers. The goals of HCI are to produce usable, safe, and functional systems. These goals can be summarized as "to develop or improve the safety, utility, effectiveness, efficiency and usability of systems that include computers." The term 'system' refers not just to hardware and software but to the entire environment—including people at work, home, or leisure. Utility refers to the functionality of a system. Usability is a key concept concerned with making systems easy to learn and easy to use.
Part of understanding user needs involves clarifying the primary objective: designing an efficient system for high productivity, a challenging system for learning, or something else. These are called usability goals and user experience goals. Usability goals meet specific criteria (e.g., efficiency), while user experience goals concern the quality of the user experience (e.g., being aesthetically pleasing).
🔑 Definition — HCI: "Human-Computer Interaction is a discipline concerned with the design, evaluation and implementation of interactive computing systems for human use and with the study of major phenomena surrounding them" (ACM/IEEE).
Usability goals
Usability involves optimizing interactions to enable users to carry out activities at work, school, and in everyday life. It is broken down into six specific goals.
Effectiveness It is a very general goal and refers to how good a system is at doing what it is supposed to do.
🔑 Definition — Effectiveness: How good a system is at doing what it is supposed to do.
Efficiency It refers to the way a system supports users in carrying out their tasks.
🔑 Definition — Efficiency: The way a system supports users in carrying out their tasks.
Safety It involves protecting users from dangerous conditions and undesirable situations. This includes two aspects: first, external working conditions (e.g., controlling X-ray machines remotely); second, helping users avoid accidental unwanted actions. Safety mechanisms include: preventing serious errors by not placing quit/delete commands next to save commands, and providing means of recovery should errors occur. Undo facilities and confirmatory dialog boxes (e.g., "Are you sure you want to delete all these messages?") are common safety features.
🔑 Definition — Safety: Protecting users from dangerous conditions and undesirable situations, including preventing errors and providing recovery mechanisms.
Utility It refers to the extent to which the system provides the right kind of functionality so users can do what they need or want to do. An example of high utility is an accounting software package providing powerful computational tools. An example of low utility is a drawing tool that forces users to use only polygon shapes and does not allow freehand drawing.
🔑 Definition — Utility: The extent to which a system provides the right kind of functionality for users' needs.
Learnability It refers to how easy a system is to learn to use. People want to get started straight away without spending too much time learning. This is especially important for everyday products (e.g., interactive TV, email) and infrequently used systems. For complex systems (e.g., web authoring tools), people are prepared to spend longer learning. A key concern is determining how much time users are willing to spend learning.
🔑 Definition — Learnability: How easy a system is to learn to use.
Memorability It refers to how easy a system is to remember how to use, once learned. This is especially important for infrequently used systems. Users should be able to remember or be rapidly reminded how to use a system without having to relearn tasks. Design supports memorability through meaningful icons, command names, menu options, and structuring options into relevant categories.
🔑 Definition — Memorability: How easy a system is to remember how to use, once learned.
"Don't Make me THINK, is the key to a usable product"
User experience goals
The emergence of new technologies (virtual reality, web, mobile computing) in diverse application areas (entertainment, education, home) has brought about wider concerns. Beyond improving efficiency and productivity, interaction design increasingly concerns itself with creating systems that are:
- Satisfying
- Enjoyable
- Fun
- Entertaining
- Helpful
- Motivating
- Aesthetically pleasing
- Supportive of creativity
- Rewarding
- Emotionally fulfilling
These goals are concerned primarily with the user experience—what the interaction with the system feels like to the users. User experience goals differ from objective usability goals because they are concerned with how users experience a product from their perspective, rather than assessing how useful or productive a system is. Recognizing trade-offs between usability and user experience goals is important, as some combinations may be incompatible (e.g., designing a process control system that is both safe and fun).
💡 Why this matters: Understanding the difference between usability and user experience goals helps designers make conscious trade-offs based on user needs. A banking app prioritizes safety and efficiency, while a game prioritizes fun and enjoyment.
4.2 Evolution of HCI
HCI takes place within a social and organizational context. Knowledge of human psychological and physiological abilities and limitations is important, including human information processing, language, communication, interaction, and ergonomics. Equally essential is knowledge of computer hardware and software possibilities, including input techniques, dialogue technique, dialogue genre/style, computer graphics, and dialogue architecture. This knowledge is brought together in the design and development of systems. Evolution plays an important role by enabling designers to check that their ideas match user needs.
Three landmark systems along this evolutionary path are the Dynabook, the Star, and the Apple Lisa. A unifying theme is that they provided effective, easy interaction for both novices and experts, with visual-spatial interfaces where objects could be directly manipulated with immediate feedback.
Dynabook Alan Kay designed the first object-oriented programming language, Smalltalk, in the 1970s, which became the basis for windows technology. Kay envisioned the Dynabook—a notebook-sized computer with a keyboard on the bottom and a high-resolution screen at the top.
Star The Xerox Star drew on PARC's ideas to create an integrated system for office environments, designed for users with no previous computer experience. The Star pioneered icons, moveable scrollable windows, and intermixed text and graphic images. The graphic user interfaces (GUIs) of today are variants of this original design.
Key design principles of the Star included:
Direct manipulation The core concept used a bitmapped screen to present direct visual representations of objects. Using a desktop metaphor, documents, printers, folders, and other office objects were depicted on screen. To print a document, the user could point to the document icon and printer icon while using a key to indicate a Copy operation.
WYSIWYG (what you see is what you get) Unlike previous programs where users created a programming-like representation and compiled it, the user works directly with the desired form through direct manipulation. The Star user could intermix text, tables, graphs, drawings, and mathematical formulas.
Consistency of commands Because a single development group developed all Star applications, a coherent and consistent design language was possible. The Star keyboard embodied generic commands used consistently across all applications: Move, Copy, Delete, Open, Show Properties, and Same. Evoking a command produced the same behavior regardless of the object type. Through property sheets, users could manipulate element-specific aspects.
Other unique features included attention to the communicative aspects of graphic design, integration of an end-user scripting language (CUSP), and mechanisms for internationalization.
Lisa by Apple The Apple Lisa introduced the GUI that started it all—with a mouse and pull-down menus. The first Lisa had dual 5.25 inch floppy drives and an external hard drive. The Lisa 2/10 moved the hard drive inside, lost one floppy drive, and shared the new 3.5-inch floppy with the Macintosh. Lisa was later marketed as the Macintosh XL, lacking the ROM toolbox built into every Macintosh, so it used MacWorks to emulate a Macintosh.
Earlier groundwork included Licklider (1960), who visualized a symbiotic relationship between humans and computers, and Sutherland (1963), who developed the Sketchpad system at MIT, introducing the ability to display, manipulate, and copy pictures on screen and the use of the light pen.
Alongside graphic interfaces, interactive text processing systems evolved. The underlying philosophy of WYSIWYG ("what you see is what you get") displayed documents on screen exactly as they would look in printed form, contrasting with earlier editors where commands were embedded in text.
Interesting transatlantic differences emerged: USA pioneers were concerned with how computers could enrich lives and facilitate creativity. European researchers (1980s) focused on constructing theories of HCI and developing design methods ensuring user needs were considered. One major European contribution was formalizing the concept of usability (Shackel, 1981).
During the 1970s, the notion of user interface (also called Man-Machine Interface or MMI) became a general concern. Moran defined this as "those aspects of the system that the user comes in contact with," meaning "an input language for the user, an output language for the machine, and a protocol for interaction."
Academic researchers focused on the capabilities and limitations of human users—understanding the 'people side' of interaction. As the field developed, it became clear that other aspects—management, organizational issues, health hazards—all contribute to the success or failure of computer systems.
⭐ Key Takeaways
The six usability goals (effectiveness, efficiency, safety, utility, learnability, memorability) are objective criteria for evaluating how well a system supports user tasks, while user experience goals (enjoyment, fun, aesthetic appeal, etc.) capture subjective feelings about interaction. Designers must recognize trade-offs between these two categories, as some goals may be incompatible for certain systems. The evolution of HCI is marked by landmark systems—Dynabook (vision of portable computing), Star (pioneering direct manipulation, WYSIWYG, and consistency), and Lisa (popularizing the GUI)—which established foundational design principles still used today. Understanding both human capabilities/limitations and technological possibilities is essential for creating effective interactive systems.
🧠 Quick Revision Questions
- What are the six specific usability goals, and how does each differ in focus?
- How do user experience goals differ from usability goals in terms of what they measure?
- What were the three key design principles of the Xerox Star, and how did each improve user interaction?
- Who developed the Dynabook concept, and what was its significance for personal computing?
- What was the difference between US and European research approaches to HCI during its early evolution?
📘 Lecture 5 — Discipline of Human Computer Interaction
📖 Overview: This lecture introduces Human-Computer Interaction (HCI) as a bridging discipline between human behavior and technology. It explores the concept of quality beyond mere conformance to specifications, and examines the interdisciplinary roots of HCI. The lecture also presents a detailed case study of a ticketing system to illustrate how HCI factors interact in real-world design.
🗂️ Topics Covered
This lecture covers the definition of HCI as a bridging discipline, the concept of quality from various perspectives including its relationship to usability, and the interdisciplinary nature of HCI. Key topics include the eight categories of factors in HCI design (organizational, environmental, health and safety, the user, cognitive processes, comfort, user interface, and task/system constraints), a detailed case study of a travel agency ticketing system, and the various contributing disciplines to HCI such as cognitive psychology, social psychology, ergonomics, linguistics, philosophy, sociology, anthropology, artificial intelligence, computer science, and engineering/design.
📝 Lecture Summary
5.1 Quality
The lecture begins by defining quality from several perspectives. According to the American Heritage Dictionary, quality is a “characteristic or attribute of something,” referring to measurable characteristics. The British Defense Industries Quality Assurance Panel defines quality as “conformance to specifications,” measuring how well design specifications are followed during manufacturing. Philip Crosby similarly states, “Quality is conformance to requirements,” where lack of conformance means lack of quality. Juran defines it as “fitness for purpose or use,” while Edward Deming says, “Quality is a predictable degree of uniformity and dependability, at low cost and suited to the market.” R J Mortiboys equates quality with “customer needs and expectations,” and Mike Robinson defines it as “meeting the (stated) requirements of the customer- now and in the future.” Armand Feigenbaum defines it as the total composite product and service characteristics that meet customer expectations in use. A final definition states, “Totality of characteristics of an entity that bear on its ability to satisfy stated and implied needs.”
From an HCI perspective, the lecture argues that quality is beyond meeting specifications or requirements. If the specifications themselves are incomplete, the product may still lack quality. True quality is measured by the end user's expectations and needs. The more usable a product is for the end user, the higher its quality. This connects quality directly to usability. The lecture defines software quality as the extent to which a software product exhibits these characteristics:
- Functionality
- Reliability
- Usability
- Efficiency
- Maintainability
- Portability
🔑 Definition — Conformance to specifications: The measure of the degree to which design specifications are followed during manufacturing. 💡 Why this matters: This lecture challenges the traditional view of quality, arguing that for HCI, quality is not just about meeting a written spec but about how well the product serves its actual end users. A product that meets all specs but is unusable is not a quality product.
5.2 Interdisciplinary nature of HCI
HCI is a discipline that is neither the study of humans nor technology alone, but the bridging between the two. A designer must always consider both what technology can do and what people will do with it. Failing to consider either leads to poor design. The main factors in HCI design are illustrated in a figure and include:
- Organizational Factors: Training, job design, politics, roles, work organization.
- Environmental Factors: Noise, heating, ventilation, lighting.
- Health and Safety: Stress, headaches, musculo-skeletal disorders.
- The User: Motivation, enjoyment, satisfaction, personality, experience level.
- Cognitive Processes and Capabilities
- Comfort Level: Seating, equipment layout.
- User Interface: Input devices, output displays, dialogue structures, use of color, icons, commands, graphics, natural language.
- Task Factors: Easy, complex, novel, task allocation, repetitive, monitoring, skills.
- System Functionality: Hardware, software, application.
- Constraints
- Productivity Factors: Increase output, increase quality, decrease costs, decrease errors, decrease labor requirements, decrease production time, increase creative ideas.
These factors interact with each other. For example, improving productivity may negatively affect user motivation if job design is ignored.
Case Study – Ticketing System: A small travel agency with multiple shops wants a new ticketing system. The current practice is slow: sales staff call airlines to check seats, get customer approval, then hand-write tickets, receipts, and itineraries. Phone connections are often busy, causing delays. Accounting is also done by hand every two weeks. Before deciding, a manager visits another company with a computerized system and finds problems: staff don’t trust the computer, can’t understand error messages, and wish to return to the old system. Sales figures are down, and staff have left. Consultants then examine user needs and company goals, recommending a system with:
- Immediate ticket booking via computer connection (solving the phone problem).
- Automatic print-out of tickets, itineraries, and receipts (reducing errors and speeding up the process).
- Direct connection between booking and accounting (speeding up accounting).
- Elimination of booking forms (reducing overheads).
The interface is designed to mimic the non-computerized task using menus and forms. The consultants are optimistic about customer satisfaction. However, they also note that for success, the layout must be changed for comfort, staff need training, job design must be changed (e.g., support during change, handling computer malfunctions), employment conditions must be examined (e.g., rewards for more transactions), and staff relations (e.g., feelings of elitism) must be managed.
HCI understands the complex relationship between humans and computers, which are two distinct “species.” Successful integration requires understanding both. HCI borrows from disciplines concerned with both:
- Human: Cognitive Psychology, Social/Organizational Psychology, Ergonomics/Human Factors, Linguistics, Philosophy, Sociology, Anthropology.
- Machine: Computer Science, Artificial Intelligence.
- Other: Engineering, Design.
Cognitive Psychology is concerned with understanding human behavior and mental processes through the notion of information processing. It characterizes these processes in terms of their capabilities and limitations.
Social and Organizational Psychology studies human behavior in a social context. Its four core concerns are: the influence of one individual on another, the impact of a group on its members, the impact of a member on a group, and the relationship between groups. Its role is to inform designers about social and organizational structures and how computers influence working practices.
Ergonomics or Human Factor developed to define and design tools and artifacts for various environments to suit user capabilities and capacities. The ergonomist’s role is to translate scientific information into design to maximize safety, efficiency, reliability, make tasks easier, and increase comfort and satisfaction.
Linguistics is the scientific study of language. It helps HCI understand issues like the order of words in command languages (e.g., delete ‘xyz’ vs. ‘xyz’ delete).
Philosophy, Sociology, and Anthropology consider the implication of IT on society. They apply social science methods to design and evaluation, providing a more accurate description of user interaction. An application is Computer Supported Cooperative Work (CSCW) , which designs shared software and hardware for groups.
Artificial Intelligence (AI) is concerned with designing intelligent computer programs that simulate intelligent human behavior. Its relationship to HCI involves user needs when interacting with an intelligent interface, such as using natural language or speech and having the system explain its advice.
Computer Science provides knowledge about technology capabilities and ideas for harnessing this potential. It develops techniques for software design, development, and maintenance.
Engineering and Design contribute through model building, empirical testing, and creative skills. Engineering’s main influence on HCI is through software engineering. Design, particularly graphic design, is a well-established discipline with benefits for HCI.
🔑 Definition — Ergonomics: A discipline that defines and designs tools and artifacts for different environments to suit the capabilities and capacities of users, with the objective of maximizing safety, efficiency, and comfort. 📐 Formula: (No formula) 📌 Example: In the ticketing system case study, ergonomics would be applied to ensure the layout of the agency is changed to make it comfortable for staff to operate the computer while still allowing contact with customers.
⭐ Key Takeaways
A student must remember that HCI is fundamentally a bridging discipline, requiring simultaneous consideration of human capabilities and technological possibilities. The concept of quality in HCI goes beyond simple conformance to specifications, placing the end user's usability and needs as the ultimate measure. The numerous factors in HCI design (organizational, environmental, user, task, etc.) are highly interdependent, and changes in one area can have ripple effects on others, as shown in the ticketing system case study. Finally, HCI is inherently interdisciplinary, drawing critical knowledge from fields like cognitive psychology, ergonomics, computer science, and social sciences to create successful human-computer interactions.
🧠 Quick Revision Questions
- According to the lecture, why is the traditional definition of quality as "conformance to specifications" insufficient for HCI?
- List at least four of the main factors that must be considered in HCI design, as shown in the lecture's figure.
- In the ticketing system case study, what were three key recommendations made by the consultants to improve the system beyond just the software itself?
- Name three disciplines that contribute to the "Human" side of HCI and briefly state what each contributes.
- What is the primary role of an ergonomist, as described in this lecture?
📘 Lecture 6 — Cognitive Frameworks
📖 Overview: This lecture introduces cognitive frameworks in Human-Computer Interaction (HCI), exploring how understanding human cognition—including perception, memory, attention, and reasoning—can inform the design of interfaces. It explains classical models like information processing and GOMS, as well as modern approaches like distributed and external cognition, to help designers create more user-friendly systems.
🗂️ Topics Covered
The lecture begins by defining cognition and its importance, then covers two general modes: experiential and reflective cognition. It details the human information processing model and its extension with memory and attention, followed by the Human Processor Model and the GOMS framework (Goals, Operators, Methods, Selection Rules). Recent developments in cognitive psychology are explored, including computational and connectionist approaches. Finally, the lecture discusses External Cognition and Distributed Cognition as more situated frameworks for understanding human-computer interaction.
📝 Lecture Summary
Introduction
The lecture begins by asking you to imagine driving a car using only a keyboard (e.g., arrow keys for steering, space bar for braking). This scenario highlights how early video games required arbitrary key combinations, while modern consoles use joysticks and steering wheels that map better to human physical and cognitive capabilities. The core goal is to understand human limitations to design systems that ease interaction. Cognitive Psychology is concerned with understanding human behavior and the underlying mental processes, adopting an information processing metaphor where everything we sense is considered information. A major focus in the 1960s and 1970s was identifying the amount of information that could be processed and remembered at one time.
Cognition
Cognition is defined as “what goes on in our heads when we carry out our everyday activities.” It refers to the processes by which we gain knowledge, including understanding, remembering, reasoning, attending, and acquiring skills. The main objective in HCI has been to understand how knowledge is transmitted between humans and computers. Specific cognitive processes include:
- Attention
- Perception and recognition
- Memory
- Learning
- Reading, speaking, and listening
- Problem solving, planning, reasoning, decision-making
These processes are interdependent; for example, learning for an exam requires attending, perceiving, reading, thinking, and remembering.
6.2 Modes of Cognition
Norman (1993) distinguishes between two general modes:
- Experiential cognition: The state of mind in which we perceive, act, and react to events around us effectively and effortlessly. It requires a certain level of expertise and engagement. Examples include driving a car, reading a book, and playing a video game.
- Reflective cognition: Involves thinking, comparing, and decision-making. This leads to new ideas and creativity. Examples include designing, learning, and writing a book.
💡 Why this matters: Both modes are essential but require different kinds of technological support. For instance, a GPS supports experiential cognition while a spreadsheet supports reflective cognition.
Information processing models the mind as an information processor. It is thought to operate through four ordered, unidirectional, and sequential stages:
- Encoding: Information from the environment is transformed into an internal representation.
- Comparison: The internal representation is compared with memorized representations.
- Response Selection: A decision is made on a response to the encoded stimulus.
- Response Execution: The response is organized and an action is performed.
📐 Formula: (Stages of Information Processing) Input → Stage 1 (Encoding) → Stage 2 (Comparison) → Stage 3 (Response Selection) → Stage 4 (Response Execution) → Output
📌 Example: To find a friend’s phone number, you must: identify the words, retrieve their meaning, understand the sentence, search memory for the number, generate a plan, and formulate a verbal answer.
Extending the model involves including attention (which selects information to process) and memory (which stores information for later use).
6.3 Human Processor Model
This model is based on the information-processing view and models cognition as a series of processing stages involving perceptual, cognitive, and motor processors. These processors are organized in relation to one another to predict user performance (e.g., reaction time) and identify bottlenecks when overloaded.
🔑 Definition — Human Processor Model: A model of the cognitive processes of a user interacting with a computer, which predicts how long a user will take to carry out various tasks, enabling comparisons between different interfaces (e.g., different word processors). It is based on the idea that cognitive activities involve people interacting with external representations (books, computers, environmental cues) rather than just internal processes.
6.4 GOMS
Card et al. developed the GOMS (Goals, Operators, Methods, and Selection Rules) family of models to translate qualitative descriptions of user performance into quantitative measures.
- Goals: Describes what the user wants to achieve. They represent a ‘memory point’ from which the user can evaluate what to do and to which they may return if errors occur.
- Operators: The lowest level of analysis; basic actions the user must perform (e.g., press a key, read a dialog box). Granularity is flexible (e.g., “issue command” vs. “move mouse to menu bar”).
- Methods: Ways a goal can be split into sub-goals; there are typically several ways to achieve a goal.
- Selection: A means of choosing between competing methods when multiple methods exist for the same goal.
Recent focus has been on Knowledge Representation Models (how knowledge is represented), Mental Models (representations people construct of themselves, others, and the environment to guide behavior), User Interaction Learning Models (how users learn to interact), Conceptual Models (ways systems are understood by different people), and Interface Metaphors (GUIs that are electronic counterparts to physical objects).
6.5 Recent Development in Cognitive Psychology
Since the 1980s, there has been a move away from the strict information-processing framework towards other theoretical approaches.
- Computational Approaches: Adopt the computer metaphor as a theoretical framework, but focus on modeling what is involved when information is processed (e.g., goals, planning, action) rather than when and how much. They analyze how information is organized, classified, and retrieved.
- Connectionist Approaches (Neural Networks / Parallel Distributed Processing): Reject the computer metaphor and instead adopt the brain metaphor. Cognition is represented at the level of neural networks of interconnected nodes. All cognitive processes are viewed as activations of nodes and connections between them.
6.6 External Cognition
External Cognition explains the cognitive processes involved when we interact with different external representations. Its main goal is to explicate the cognitive benefits of using different representations.
The main benefits include:
- Externalizing to reduce memory load: Transforming knowledge into external representations (e.g., writing down birthdays, appointments) to remember to do something (e.g., buy a card) or when to do something (e.g., send it by a certain date).
- Computational offloading: Using a tool or device in conjunction with an external representation to help carry out a computation (e.g., using pen and paper to solve a math problem).
- Annotating and cognitive tracing:
- Annotating: Modifying external representations, such as crossing off or underlining items in a to-do list.
- Cognitive tracing: Externally manipulating items into different orders or structures (e.g., creating different piles of work).
Information Visualization is a design principle based on external cognition: providing external representations at the interface that reduce memory load and facilitate computational offloading, thereby extending or amplifying cognition (e.g., financial forecasting tools, identifying programming bugs).
6.7 Distributed Cognition
Distributed Cognition is a theoretical framework that goes beyond the individual to conceptualize cognitive activities as embodied and situated within the work context. It describes cognition as distributed across individuals, computer systems, and other technology in a setting, which together are called a functional system. Examples include ship navigation and air traffic control.
The main goal is to analyze how the different components of the functional system are coordinated, focusing on how information is propagated through the system in terms of technological, cognitive, social, and organizational aspects. It analyzes how information moves and transforms between different representational states.
⭐ Key Takeaways
The most critical takeaway is that understanding human cognitive processes—attention, perception, memory, and learning—is fundamental to designing effective and intuitive interfaces. Classical models like the information processing model and GOMS provide frameworks for predicting user performance and comparing interface efficiency, but they are limited by their focus on internal, individual cognition. Modern frameworks like external cognition and distributed cognition offer more complete explanations by situating cognitive activity in the real world, emphasizing how people use external tools and collaborate with others. Ultimately, designers must recognize the interdependence of cognitive processes and support both experiential (effortless) and reflective (analytical) modes of thought.
🧠 Quick Revision Questions
- Define cognition and explain why it is a central concept in HCI.
- Distinguish between experiential and reflective cognition, giving an example of each and explaining why each requires different technological support.
- Briefly describe the four stages of the human information processing model. What is a major limitation of this model according to the lecture?
- What does GOMS stand for? Explain the role of each component and how this model is used to evaluate user interfaces.
- What is external cognition, and what are three key cognitive benefits it provides? Give an example of how you use each benefit in your daily life.
📘 Lecture 7 — Human Input-Output Channels – Part I
📖 Overview: This lecture introduces the study of human input-output channels in Human Computer Interaction. It covers the physiological and perceptual aspects of human vision, including the anatomy of the eye and how visual perception works. Understanding these human capabilities and limitations is essential for designing effective computer interfaces.
🗂️ Topics Covered
The lecture covers input-output channels in human-computer interaction, focusing on the five senses and effectors involved. It then delves deeply into vision, examining the human eye's physiology including rods, cones, fovea, and the blind spot, followed by visual perception covering size and depth perception, brightness perception, and color perception, along with the capabilities and limitations of visual processing including optical illusions.
📝 Lecture Summary
7.1 Input Output channels
A person's interaction with the outside world occurs through information being received (input) and sent (output). In computer interaction, the user's output becomes the computer's input and vice versa. The five major senses are sight, hearing, touch, taste, and smell — of which the first three are most important to HCI. Taste and smell do not currently play a significant role in general computer systems.
The human effectors include limbs, fingers, eyes, head, and vocal system. In computer interaction, fingers play the primary role through typing or mouse control. For example, using a personal computer with a mouse and keyboard involves receiving information primarily by sight, with hearing and touch also providing feedback. The user sends information using hands through keystrokes or mouse movements.
7.2 Vision
Human vision is a highly complex activity with physical and perceptual limitations, yet it is the primary source of information. Visual perception can be divided into two stages: the physical reception of the stimulus from the outside world, and the processing and interpretation of that stimulus.
The human eye
Vision begins with light. Light is reflected from objects and focused upside down on the back of the eye. The cornea and lens at the front of the eye focus light into a sharp image on the retina. The retina is light sensitive and contains two types of photoreceptor: rods and cones.
🔑 Definition — Rods: Highly sensitive to light, allowing us to see under low illumination, but unable to resolve fine detail and subject to light saturation. There are approximately 120 million rods per eye, mainly situated towards the edges of the retina, dominating peripheral vision.
📌 Example — Rod saturation: When moving from a darkened room into sunlight, rods that have been active become saturated by the sudden light, causing temporary blindness. Cones do not operate either as they are suppressed by the rods.
🔑 Definition — Cones: Less sensitive to light than rods and can tolerate more light. There are three types of cone, each sensitive to a different wavelength of light, allowing color vision. The eye has approximately 6 million cones, mainly concentrated on the fovea.
🔑 Definition — Fovea: A small area of the retina on which images are fixated.
🔑 Definition — Blind spot: An area of the retina where the optic nerve enters the eye, containing no rods or cones. The visual system compensates for this so we are normally unaware of it.
The retina also has specialized nerve cells called ganglion cells:
- X-cells: Concentrated in the fovea, responsible for early detection of pattern.
- Y-cells: More widely distributed, responsible for early detection of movement. This means we can perceive movement in peripheral vision even if we cannot detect pattern changes there.
7.3 Visual perception
Visual perception involves more than just the physical mechanism of vision. The information received by the visual apparatus must be filtered and passed to processing elements that allow us to recognize coherent scenes, disambiguate relative distances, and differentiate color.
Perceiving size and depth
The visual system easily interprets images to account for size and distance. The size of an image on the retina is specified as visual angle, which is affected by both the size of the object and its distance from the eye.
🔑 Definition — Visual angle: The angle between two lines drawn from the top and bottom of an object to a central point on the front of the eye. It indicates how much of the field of view is taken by the object. Visual angle is measured in degrees or minutes of arc (1 degree = 60 minutes of arc, 1 minute = 60 seconds of arc).
📌 Example — Visual angle: If two objects are at the same distance, the larger one will have the larger visual angle. If two objects of the same size are placed at different distances, the furthest one will have the smaller visual angle.
🔑 Definition — Visual acuity: The ability of a person to perceive fine detail. A person with normal vision can detect a single line if it has a visual angle of 0.5 seconds of arc. Spaces between lines can be detected at 30 seconds to 1 minute of visual arc.
🔑 Definition — Law of size constancy: Our perception of an object's size remains constant even if its visual angle changes. A person's height is perceived as constant even if they move further away. This indicates that our perception of size relies on factors other than visual angle, such as depth perception.
Cues for depth perception:
- Overlap: If objects overlap, the partially covered object is perceived as being further away.
- Size and height in the field of view provides a cue to distance.
- Familiarity: If we expect an object to be a certain size, we can judge its distance accordingly.
Perceiving brightness
🔑 Definition — Brightness: A subjective reaction to the level of light. It is affected by luminance, which is the amount of light emitted by an object. Luminance depends on the amount of light falling on the object's surface and its reflective properties.
🔑 Definition — Contrast: A function of the luminance of an object and the luminance of its background.
In dim lighting, rods predominate vision. Since there are fewer rods on the fovea, objects in low lighting are more visible in peripheral vision. In normal lighting, cones take over.
Visual acuity increases with increased luminance, suggesting high display luminance is beneficial. However, as luminance increases, flicker also increases. The eye perceives a light switched on and off rapidly as constantly on if the switching speed is less than 50 Hz. Flicker is more noticeable in peripheral vision, meaning larger displays appear to flicker more.
💡 Why this matters: This trade-off between luminance and flicker directly affects display design — designers must balance clarity with visual comfort.
Perceiving color
Color is usually regarded as being made up of three components:
- Hue
- Intensity
- Saturation
🔑 Definition — Hue: Determined by the spectral wavelength of the light. Blues have short wavelength, greens medium, and reds long. Approximately 150 different hues can be discriminated by the average person.
🔑 Definition — Intensity: The brightness of the color.
🔑 Definition — Saturation: The amount of whiteness in the color.
By varying intensity and saturation, we can perceive approximately 7 million different colors. However, the number of colors identifiable without training is far fewer.
The eye perceives color because cones are sensitive to different wavelengths (blue, green, red). Color vision is best in the fovea and worst at the periphery. Only 3-4% of the fovea is occupied by cones sensitive to blue light, making blue acuity lower.
🔑 Definition — Color blindness: Around 8% of males and 1% of females suffer from color blindness, most commonly being unable to discriminate between red and green.
The capabilities and limitations of visual processing
Visual processing involves the transformation and interpretation of a complete image from light on the retina. Our expectations affect image perception — for example, if we know an object's size, we perceive it as that size regardless of distance.
Visual processing compensates for:
- Movement of the image on the retina — the perceived image remains stable even as we move
- Color and brightness of objects — perceived as constant despite luminance changes
This ability to interpret and exploit expectations can resolve ambiguity, as context allows disambiguation (e.g., interpreting a shape as "B" or "13" based on surrounding characters).
However, expectations can also create optical illusions:
- Muller-Lyer illusion: Lines of equal length appear different due to arrowhead direction
- Ponzo illusion: Lines of equal length appear different due to perspective cues
- Proofreading illusion: Our expectations cause us to miss errors in familiar text
Our perception of geometric shapes is not exact — we tend to magnify horizontal lines and reduce vertical lines, so a square needs to be slightly increased in height to appear square. Lines appear thicker if horizontal rather than vertical. The optical center of a page is perceived as slightly above the actual center, so symmetrical arrangement around the actual center appears too low.
⭐ Key Takeaways
The most critical knowledge from this lecture is understanding that human vision involves two distinct stages—physical reception of light through the eye's anatomy and cognitive processing and interpretation—both of which impose limitations on interface design. Students must remember that rods enable peripheral and low-light vision while cones enable color vision with concentration in the fovea, and that visual angle determines object perception but is modified by size constancy laws. Visual perception of brightness involves luminance and flicker thresholds (50 Hz), while color perception involves hue, intensity, and saturation with important implications given color blindness affects 8% of males. Finally, optical illusions demonstrate that visual processing compensates for and interprets images based on expectations, which designers must account for in layout and visual element placement.
🧠 Quick Revision Questions
-
What are the two stages of visual perception, and why must interface designers understand both?
-
Explain the functional differences between rods and cones, including their approximate numbers per eye and where each is concentrated on the retina.
-
What is visual angle, and how does it relate to the law of size constancy?
-
What is the flicker fusion threshold for the human eye, and why does it matter for display design?
-
What are the three components of color perception, and what percentage of males suffer from red-green color blindness?
📘 Lecture 8 — Human Input-Output Channels Part II
📖 Overview: This lecture continues the exploration of human input-output channels, focusing on how humans perceive and process visual information through color, depth, and movement. It also covers hearing, touch (haptic perception), and the mechanics of reading and movement, all of which are critical for designing effective human-computer interfaces.
🗂️ Topics Covered
The lecture covers color theory including the color wheel, primary, secondary, and tertiary colors, color harmony, and color context. It also delves into stereopsis (3D vision), the reading process, the structure and function of hearing, haptic perception, and movement including reaction time, movement time, and various motion perception phenomena.
📝 Lecture Summary
8.1 Color Theory
Color theory encompasses definitions, concepts, and design applications. The color wheel is a circular diagram of colors, first developed by Sir Isaac Newton in 1666, based on red, yellow, and blue. Primary colors are red, yellow, and blue—the three pigment colors that cannot be mixed from other colors. Secondary colors are green, orange, and purple, formed by mixing primary colors. Tertiary colors are yellow-orange, red-orange, red-purple, blue-purple, blue-green, and yellow-green, formed by mixing one primary and one secondary color.
Color harmony is a pleasing arrangement of parts that engages the viewer and creates balance. Extreme unity leads to under-stimulation; extreme complexity leads to over-stimulation. Harmony is a dynamic equilibrium.
Analogous colors are any three colors side by side on a 12-part color wheel (e.g., yellow-green, yellow, yellow-orange). Complementary colors are any two colors directly opposite each other (e.g., red and green), creating maximum contrast and stability. Natural harmony comes from nature, like red, yellow, and green together.
Color context refers to how color behaves relative to other colors. For example, red appears more brilliant against black, duller against white, and lifeless against orange. As we age, the lens becomes yellow and absorbs shorter wavelengths, so blue should not be used for text. Older people need brighter colors.
💡 Why this matters: Color choices directly affect readability, user fatigue, and accessibility in interfaces.
Guidelines include:
- Opponent colors (red & green, yellow & blue) go well together
- Pick non-adjacent colors on the hue circle
- It's hard to detect changes in reds, purples, & greens; easier in yellows & blue-greens
- Older users need higher brightness levels
- Use both brightness & color differences for edges
- Avoid red & green in periphery; use yellows & blues
- Avoid pure blue for text, lines, & small shapes
- Blue makes a fine background color
- Avoid adjacent colors that differ only in blue
- Mixtures of colors should differ in 2 or 3 colors
- Accurate color discrimination is at ±60 degrees of straight head position
- Limit of color awareness is ±90 degrees
8.2 Stereopsis
Stereopsis is the direct sensing of an object's distance by comparing images received by the two eyes. It exists in animals with overlapping optical fields. It is the most reliable depth clue and overrides all others. The pair of views is called a stereopair or stereogram. The word comes from Greek στερεοs (solid) and οψιs (appearance). Haplopia is single vision; diplopia is double vision.
Stereopsis yields benefits for close work like fighting for cats and hand work for humans. The brain fuses the two images from the eyes. Visual perception uses many distance clues: strong clues include apparent sizes, overlapping, parallax, shadows, and perspective; weaker clues include atmospheric perspective, speed, and detail.
The interpretation is entirely mental and must be learned. Free fusion is the skill of fusing two pictures side by side by voluntarily diverging the eyes. A stereoscope helps achieve fusion without practice. When images are too different, rivalry occurs—one image is favored or a patchwork is seen. When everything corresponds except illumination or color, the fused image exhibits lustre.
The fundamentals were discovered by Charles Wheatstone in 1836. An anaglyph is a stereopair drawn in two colors viewed through colored filters. About 4% of people have defective stereopsis.
The autostereogram is a single figure that gives stereoscopic images to the two eyes (e.g., the "wallpaper illusion" discovered by H. Meyer in 1842). A random-dot auto stereogram (Julesz, 1960) requires free fusion. No image is seen until fusion occurs.
🔑 Definition — Stereopsis: The direct sensing of distance by comparing images from two eyes. 📌 Example: Viewing a waterfall for a minute, then looking at a stationary object causes the object to appear to move upward (motion after effect).
8.3 Reading
Reading has several stages: perceiving the visual pattern of words, decoding with reference to internal language representation, and syntactic/semantic analysis. During reading, the eye makes jerky movements called saccades followed by fixations. Perception occurs during fixations, accounting for about 94% of time. Regressions are backward eye movements; complex text causes more regressions.
8.4 Hearing
Hearing begins with sound waves. The ear has three sections: outer ear, middle ear, and inner ear. The outer ear includes the pinna (visible structure) and auditory canal (which protects the middle ear and amplifies sounds). The middle ear is connected to the outer ear by the tympanic membrane (eardrum) and contains the ossicles (smallest bones in the body). The inner ear contains the cochlea (liquid-filled) with delicate hair cells or cilia that release a chemical transmitter causing impulses in the auditory nerve.
Pitch is the frequency of sound (low frequency = low pitch). Loudness is proportional to amplitude (frequency constant). Timbre relates to the type of sound.
🔑 Definition — Pitch: The frequency of a sound wave. 📐 Formula: Low frequency = Low pitch; High frequency = High pitch 📌 Example: A violin and a piano playing the same note at the same loudness have different timbres.
Sound characteristics: Audible range is 20 Hz to 15 KHz. The ear can distinguish changes less than 1.5 Hz but is less accurate at higher frequencies. The Cocktail Party Effect is the auditory system's ability to filter sounds.
8.5 Touch
Haptic perception involves sensors in the skin, hand, and arm. It includes mechanoreceptors (deformation, thermo reception, vibration), plus receptors in muscles, tendons, and joints. This combination of kinesthetic and sensory perception creates strong neural pathways.
Haptics vs. Vision: Vision is rapid and holistic, superior for macro geometry (shape). Haptics is superior for micro geometry (texture), detecting properties like roughness, hardness, wetness, stickiness, and temperature. Together they are superior for many learning contexts.
🔑 Definition — Haptic perception: The sense of touch involving sensors in the skin, hand, and arm.
8.6 Movement
A simple action like hitting a button involves: stimulus reception → brain processing → response generation → muscle response. These stages divide into reaction time and movement time.
- Reaction time: Auditory signal ~150ms, Visual signal ~200ms, Pain ~700ms
- Movement time: Depends on age and fitness
Movement perception involves different ways of following moving objects: moving eyes alone, moving head alone, or combinations. All three give similar perception.
Real movement is when the physical stimulus actually moves. Apparent movement occurs when there is no physical motion. Examples:
- Motion After Effect (MAE): After viewing motion in one direction, stationary objects appear to move in the opposite direction.
- Phi phenomenon: Two lights alternately flashing appear as a single moving light (used in movies and marquees).
- Induced motion: When a nearby vehicle moves, you feel as if you are moving.
- Auto kinetic movement: A small dim light in a dark room appears to move randomly.
The visual system has motion detectors that undergo spontaneous activity. Adaptation of these detectors causes the MAE. Electro physiologists have discovered cortical neurons specialized for movement in specific directions.
💡 Why this matters: Understanding movement perception helps design animations, notifications, and responsive interfaces.
🔑 Definition — Saccades: Jerky eye movements during reading followed by fixations. 🔑 Definition — Phi phenomenon: The perception of motion from alternating stationary lights.
⭐ Key Takeaways
Color theory is essential for interface design—use the color wheel, understand harmony (analogous, complementary, natural), and follow guidelines (avoid pure blue for text, use both brightness and color differences, consider aging users). Stereopsis (3D vision) relies on comparing two eye images and is the most reliable depth clue, overriding all others. Reading involves saccades and fixations (94% of time), with regressions for complex text. Hearing processes pitch, loudness, and timbre, with an audible range of 20 Hz to 15 KHz. Haptic perception is superior for texture and material properties, while vision dominates for shape. Movement perception includes reaction time (auditory fastest at 150ms) and phenomena like MAE, phi phenomenon, induced motion, and auto kinetic movement, all crucial for natural interface interactions.
🧠 Quick Revision Questions
- What are the three primary colors in traditional color theory, and why can't they be mixed from other colors?
- Explain the difference between analogous and complementary color schemes with examples.
- What is stereopsis, and why is it considered the most reliable depth clue?
- Describe the three main sections of the human ear and their functions in hearing.
- Compare the reaction times for auditory, visual, and pain stimuli, and list two examples of apparent movement perception.
📘 Lecture 9 — Cognitive Process - Part I
📖 Overview: This lecture introduces the first two cognitive processes from the Extended Human Processing Model: Attention and Memory. It explains how attention allows us to select and focus on relevant information, and details the structure and function of human memory, including sensory, short-term, and long-term memory stores. Understanding these processes is crucial for designing interfaces that align with human cognitive capabilities and limitations.
🗂️ Topics Covered
This lecture begins by listing the cognitive processes (Attention, Memory, Perception, Learning, Reading/Speaking/Listening, Problem-solving). It then provides a detailed explanation of Attention, including its types (focused, divided, voluntary, involuntary) and how to focus attention at the interface. The second half covers Memory, explaining the multi-store model with sensory, short-term, and long-term memory, including the recency and primacy effects, and concludes with a revised memory model.
📝 Lecture Summary
9.1 Attention
Attention is the process of selecting a specific thing to concentrate on from all the available possibilities at a given moment. As psychologist Williams James described, it involves "taking possession of the mind by one out of several simultaneously possible objects or trains of thought," requiring withdrawal from some things to deal effectively with others. Attention can involve our auditory senses (e.g., waiting for your name to be called in a dentist's office, based on pitch, timber, and intensity) or visual senses (e.g., scanning a newspaper for football results, based on color and location). The ease of attending depends on whether we have clear goals and whether the needed information is salient in the environment.
Our goals guide our attention. If we know exactly what we want (e.g., who won the World Cup), we actively search for matching information. If our goal is vague (e.g., deciding what to eat at a restaurant), we browse, allowing our attention to be drawn to salient items. Information presentation also greatly influences attention. Structuring information meaningfully makes it easier to find. For example, a screen with information grouped into vertical categories with spaces between columns was found to be searchable in an average of 3.2 seconds, while a screen with the same information bunched together took an average of 5.5 seconds to search.
🔑 Definition — Attention: The process of selecting things to concentrate on, at a point in time, from the range of possibilities available. 📐 Formula/Model: Focused Attention (attending to one event from competing stimuli) vs. Divided Attention (attending to more than one thing at a time). Attention can be Voluntary (conscious effort) or Involuntary (grabbed by salient stimuli). 📌 Example - Focused vs. Divided: Having a conversation with one person is focused attention. Having a conversation while intermittently watching someone else is divided attention. An example of involuntary attention is being distracted from your work by music from the next room.
Focusing attention at the interface
For HCI, understanding attention is key to designing effective interfaces. Designers must consider how to focus users' attention on the relevant information for a given task, how to get their attention when needed, and how to avoid unnecessary distractions.
🔑 Definition — Structuring Information: Designing the interface so information is easy to navigate by presenting the right amount of information, grouped and ordered into meaningful parts using perceptual laws of grouping. 💡 Why this matters: Proper structuring helps the user attend to their task, not the interface itself. 📌 Example - Design Considerations: Help users stay focused by avoiding unnecessary distractions. Only create distraction or use alerts when they are truly appropriate. Make information salient when it needs to be attended to by using techniques like color, ordering, spacing, underlining, sequencing, and animation. Avoid cluttering the interface; follow the example of a crisp, simple design like Google.com.
9.2 Memory
Memory is essential for everyday activities, storing factual knowledge, actions/procedures, and our sense of identity. The multi-store model of memory describes three types of memory or memory functions.
Memory Model
The multi-store model of memory consists of three types of stores:
- Sensory store: modality-specific, holds information for a very brief period (a few tenths of a second).
- Short-term memory store: holds limited information for a short period (a few seconds).
- Permanent long-term memory store: holds information indefinitely.
Sensory memory acts as a buffer for stimuli received through the senses. A sensory memory exists for each sensory channel:
- Iconic memory: for visual stimuli. Information remains for about 0.5 seconds (e.g., the persistent image of a moving sparkler).
- Echoic memory: for aural stimuli. Allows brief playback of information (e.g., realizing you knew what someone asked even after you asked them to repeat it).
- Haptic memory: for touch.
Information is passed from sensory memory to short-term memory (STM) by attention, which filters stimuli to only those of interest. This selective attention is necessary due to our limited processing capacity and explains the cocktail party phenomenon—being able to focus on one conversation but switch attention if your name is mentioned elsewhere.
🔑 Definition — Short-term memory (STM) or Working memory: Acts as a "scratch pad" for temporary recall of information, used for fleetingly needed information (e.g., intermediate steps in mental arithmetic, holding the beginning of a sentence in mind while reading the rest). 📐 Capacity and Duration: STM has a rapid access time (~70ms) but decays rapidly (~200ms). Its capacity is limited to 7±2 chunks of information (Miller's Law). This can be increased by chunking information (e.g., remembering "22 55 36 8998 30" is easier than "54988319814237"). 📌 Example - Recency Effect: In a free recall task, items at the end of the list are recalled better because they are still active in STM. The primacy effect is the better recall for items at the beginning of the list, as they have been rehearsed more and placed into LTM.
Long-term memory (LTM) is our main resource, storing all our factual and experiential knowledge. It differs from STM in having a huge, if not unlimited, capacity, a slower access time (~0.1 seconds), and very slow forgetting.
Long-term memory structure
LTM is divided into two main types:
- Episodic memory: Represents our memory of events and experiences in a serial form. It allows us to reconstruct the actual events that took place at a given period of our lives.
- Semantic memory: A structured record of facts, concepts, and skills we have acquired. Information is derived from episodic memory. Semantic memory is thought to be structured as a semantic network, where items are associated in classes and can inherit attributes from parent classes (e.g., inferring that a specific sheepdog named Shadow has four legs and a tail). Other structures like frames and scripts organize information into data structures with slots for attribute values.
🔑 Definition — Episodic memory: The memory of events and experiences in a serial form. 🔑 Definition — Semantic memory: A structured record of facts, concepts, and skills.
9.3 Revised Memory Model
According to the revised memory model, working memory is a subset of LTM. In this view, items in working memory are activated chunks from long-term memory. Activation is supplied from other linked chunks and from sensory input. This model views human performance as a series of processing stages involving perceptual, motor, and cognitive systems, each with its own memory and processes, organized in a particular way.
⭐ Key Takeaways
The most critical concepts from this lecture are the two types of cognitive processes: attention and memory. For attention, you must understand how goals and information presentation (structuring) affect what we focus on, and the difference between focused, divided, voluntary, and involuntary attention. For memory, you must know the three stores of the multi-store model: sensory memory (iconic, echoic, haptic), short-term memory (with its 7±2 capacity and rapid decay), and long-term memory (episodic and semantic, including semantic networks). Key effects like the recency and primacy effects, and the concept of chunking, are essential for understanding how information is processed and recalled, directly informing interface design.
🧠 Quick Revision Questions
- What are the two main factors that determine how easy or difficult it is to pay attention to relevant information?
- Describe the difference between focused and divided attention, and provide an example of each from the lecture.
- What is the multi-store model of memory? Name the three stores and one key characteristic of each.
- Explain Miller's Law (the 7±2 rule) and provide an example of what "chunking" means in this context.
- What is the difference between episodic memory and semantic memory in long-term memory?
📘 Lecture 10 — Cognitive Processes - Part II
📖 Overview: This lecture continues the exploration of cognitive processes in Human Computer Interaction, focusing on learning and thinking. It covers how people learn through procedural and declarative methods, the psychology of reading/speaking/listening, and the higher-order cognitive processes of problem-solving, planning, reasoning, and decision-making. Understanding these cognitive mechanisms is crucial for designing interfaces that align with natural human learning and thinking patterns.
🗂️ Topics Covered
The lecture examines learning through procedural and declarative approaches, including practical learning strategies like the "training wheels" method and dynalinking. It then covers the three modes of language processing—reading, speaking, and listening—with their design implications. The final major section explores reflective cognition through problem-solving theories (Gestalt and Problem Space Theory), three types of reasoning (deductive, inductive, abductive), analogical mapping, and skill acquisition differences between novices and experts.
📝 Lecture Summary
10.1 Learning
Learning is divided into two categories: procedural learning, which focuses on "how to" use something (e.g., how to use a computer application), and declarative learning, which focuses on facts about something (e.g., using an application to understand a topic). Jack Carroll's research shows that people find it very hard to learn by following sets of instructions in a manual. When people encounter a computer for the first time, their common reaction is fear and trepidation, contrasting sharply with the motivation people feel when learning to drive a car. This discrepancy arises because driving is taught through actual doing, while computer learning often involves overwhelming manuals.
Experienced users are also reluctant to learn new methods from manuals, preferring to continue using familiar procedures even if less effective. People prefer learning through doing, and GUI and direct manipulation interfaces support this by allowing exploratory interaction and the ability to 'undo' actions. Carroll suggested the 'training wheels' approach, which restricts novice users to basic functions and extends them as the user gains experience, making initial learning more tractable.
Dynalinking is the process of linking and manipulating multimedia representations at the interface. It helps learners understand complex relationships by connecting concrete simulations with abstract representations. For example, a multimedia simulation of a pond ecosystem showed organisms swimming, and clicking on an organism revealed what it was and what it ate. This simulation was dynalinked with abstract food web diagrams, allowing users to see how changes in one representation affected the other. Dynalinking is especially powerful for learning complex, multi-dimensional information.
🔑 Definition — Procedural Learning: Learning focused on how to do something or how to use an object. 🔑 Definition — Declarative Learning: Learning focused on facts about something. 🔑 Definition — Training Wheels Approach: Restricting novice users to basic functions and extending them as they gain experience. 🔑 Definition — Dynalinking: The process of linking and manipulating multimedia representations at the interface to show relationships. 💡 Why this matters: Understanding learning preferences helps designers create interfaces that support exploratory learning, undo capabilities, and gradual skill development rather than forcing users through linear manual instruction.
10.2 Reading, Speaking and Listening
These three forms of language processing share the property that sentence meaning remains the same regardless of the mode of conveyance. However, the ease of reading, listening, or speaking differs based on person, task, and context. Written language is permanent while listening is transient—rereading is possible but rebroadcasting spoken information is not. Reading can be quicker than speaking or listening because text can be rapidly scanned. Listening requires less cognitive effort than reading or speaking, especially for children. Written language tends to be grammatical while spoken language is often ungrammatical. Differences in ability exist, with dyslexics having difficulty understanding written words, and people who are hard of hearing or seeing being restricted in language processing.
Design implications include keeping speech-based menus and instructions short (people struggle with more than three or four options), accentuating intonation in artificial speech voices, and providing opportunities to enlarge text without affecting formatting for those who find small text difficult to read.
🔑 Definition — Dyslexics: Individuals who have difficulties understanding and recognizing written words, making it hard to write grammatical sentences and spell correctly.
10.3 Problem Solving, Planning, Reasoning and Decision-making
These cognitive processes involve reflective cognition, which includes thinking about what to do, what the options are, and what the consequences might be. They involve conscious processing, discussion with others, and use of artifacts like maps and books. The extent to which people engage in reflective cognition depends on their level of experience. Novices often act by trial and error, making mistakes and acting inefficiently, while experts have more knowledge and can select optimal strategies and think ahead about consequences.
Reasoning uses knowledge to draw conclusions or infer new information about a domain. There are three types:
- Deductive reasoning derives the logically necessary conclusion from premises. For example: "If it is Friday then she will go to work. It is Friday. Therefore she will go to work." Validity does not require truth—a valid deduction can conflict with real-world knowledge.
- Inductive reasoning generalizes from cases we have seen to infer information about unseen cases. For example, if every elephant observed has a trunk, we infer all elephants have trunks. This inference is unreliable—it can only be proved false, never proved true.
- Abductive reasoning reasons from a fact to the action or state that caused it. For example, if Sam always drives fast after drinking, and we see Sam driving fast, we may infer she has been drinking. This is also unreliable but people hold onto explanations until evidence supports an alternative.
Problem solving is the process of finding a solution to an unfamiliar task using available knowledge. The Gestalt theory views problem solving as both reproductive (drawing on previous experience) and productive (involving insight and restructuring). Reproductive problem solving can be a hindrance if a person 'fixates' on known aspects and cannot see novel interpretations. However, Gestalt theory lacks sufficient evidence about when restructuring occurs or what insight is.
The Problem Space Theory (Newell and Simon, 1970s) proposes that problem solving centers on the problem space, which comprises problem states. People use legal state transition operators to move from an initial state to a goal state. Means-ends analysis is a heuristic where the initial state is compared with the goal state, and an operator is chosen to reduce the difference. For example, moving a desk requires recognizing the location difference and creating sub-goals like making the desk light enough to carry. Problem space searching is limited by short-term memory capacity and retrieval speed.
Analogy in problem solving uses analogical mapping—mapping knowledge from a similar known domain to a new problem. In Gick and Holyoak's experiment, only 10% of subjects solved the tumor problem (needing converging low-intensity rays) without help, but 80% succeeded when given the analogous fortress story about converging small groups of soldiers.
Skill acquisition shows differences between novices and experts. Chess experiments demonstrated that masters remembered board configurations and associated good moves using larger 'chunks' than less experienced players. Masters considered no more alternatives than novices but took less time and produced better moves. Novices group problems by superficial characteristics (objects or features), while experts group by underlying conceptual similarities.
🔑 Definition — Reflective Cognition: Cognitive processes involving thinking about what to do, options, and consequences, often requiring conscious processing. 🔑 Definition — Deductive Reasoning: Deriving the logically necessary conclusion from given premises. 🔑 Definition — Inductive Reasoning: Generalizing from observed cases to infer information about unobserved cases. 🔑 Definition — Abductive Reasoning: Reasoning from a fact to the action or state that caused it. 🔑 Definition — Problem Space Theory: The theory that problem solving centers on problem states and uses legal operators to move from initial to goal states. 🔑 Definition — Means-Ends Analysis: A heuristic where the initial state is compared with the goal state and an operator is chosen to reduce the difference. 🔑 Definition — Analogical Mapping: Mapping knowledge from a similar known domain to a new problem. 📐 Formula: Deductive Reasoning: If P then Q; P is true → Therefore Q is true. 📐 Formula: Inductive Reasoning: All observed cases of X have property Y → Therefore all X have property Y (unreliable). 📐 Formula: Abductive Reasoning: Fact F is observed; Known rule: Action A causes F → Therefore A occurred (unreliable). 📌 Example (Deductive): If it is raining then the ground is dry (premise). It is raining (premise). Therefore the ground is dry—valid deduction even though it conflicts with real-world knowledge. 📌 Example (Inductive): Every elephant ever seen has a trunk → Infer all elephants have trunks. Cannot be proved true (next elephant might be trunkless). 📌 Example (Abductive): Sam drives too fast when she has been drinking (known rule). Sam is driving too fast (observed fact) → Infer Sam has been drinking (possibly wrong—she may have an emergency). 📌 Example (Problem Space - Means-Ends): Initial state: desk at north wall. Goal state: desk by window. Difference: location. Desk is heavy (constraint). Sub-goal: make desk light (remove drawers). Operator: carry desk.
⭐ Key Takeaways
Students must understand the fundamental distinction between procedural and declarative learning and why people prefer learning through doing rather than reading manuals. The "training wheels" approach and dynalinking are critical design strategies for facilitating learning at interfaces. The three modes of language processing have distinct properties that inform design choices, particularly the limitations of speech-based interfaces. The three types of reasoning (deductive, inductive, abductive) are essential for understanding how users draw conclusions from system behavior, with abduction being particularly important for interface design because users infer causation from correlation. Finally, the Problem Space Theory with means-ends analysis, analogical mapping, and the differences between novice and expert problem-solving (chunking, superficial vs. conceptual grouping) are foundational concepts for designing systems that support users at different skill levels.
🧠 Quick Revision Questions
- What are the two types of learning, and how does the "training wheels" approach support each type?
- What is dynalinking, and how did the pond ecosystem experiment demonstrate its effectiveness for learning complex concepts?
- What are the three types of reasoning? Give a real-world example of each type and explain why abductive reasoning is particularly relevant to interface design.
- According to Problem Space Theory, what is means-ends analysis, and how does it help users solve problems within the constraints of human processing limitations?
- What differences were observed between chess masters and less experienced players in terms of board memory and problem grouping strategies?
📘 Lecture 11 — The Psychology of Actions
📖 Overview: This lecture explores how people form mental models to understand and interact with systems, and how these models can lead to errors. It examines the psychology of action through the seven stages of action cycle and classifies different types of human errors, explaining why self-blame and learned helplessness occur in human-computer interaction.
🗂️ Topics Covered
The lecture covers mental models and their characteristics, the difference between images and mental models, erroneous mental models with examples like thermostats and elevator buttons, self-blame and learned helplessness, taught helplessness, the nature of human thought and explanation with case studies of Three Mile Island and Lockheed L-1011 accidents, the action cycle including execution and evaluation stages, and a classification of errors into slips and mistakes.
📝 Lecture Summary
11.1 Mental model
The concept of mental model was first developed in the early 1640s by Kenneth Craik, who proposed that thinking involves "models, or parallels reality." Donald Norman defines mental models as "the model people have of themselves, others, the environment, and the things with which they interact. People form mental models through experience, training and instruction."
Mental models are built to make predictions about external events before carrying out an action. An important observation is that they are invariably incomplete, unstable, easily confusable, and often based on superstition rather than scientific fact.
Within cognitive psychology, Johnson-Laird (1983, 1988) explicated mental models regarding their structure and function in human reasoning. He argues that mental models are either analogical representations or a combination of analogical and propositional representations. A mental model represents the relative position of objects in an analogical manner that parallels the structure of the state of objects in the world.
🔑 Definition — Mental Model: "The model people have of themselves, others, the environment, and the things with which they interact. People form mental models through experience, training and instruction." (Donald Norman)
An important difference exists between images and mental models regarding their function. Mental models are constructed when we need to make an inference or prediction, and a conscious mental simulation may be "run" from which conclusions can be deduced. An image, conversely, is a one-off representation. A simplified analogy: an image is like a frame in a movie while a mental model is more like a short snippet of a movie.
While learning and using a system, people develop knowledge of how to use the system and, to a lesser extent, how the system works. These two kinds of knowledge are often referred to as a user's mental model. The more someone learns about a system, the more their mental model develops. TV engineers have a deep mental model of how TVs work, while an average citizen has a shallow mental model of operation.
📌 Example — Thermostat Mental Models: When arriving home to a cold house with a baby, most people set the thermostat as high as possible, believing this increases the rate of warming. This is incorrect. Two commonly held folk theories exist:
- Timer theory: The thermostat controls the proportion of time the device stays on (set higher = more time on)
- Valve theory: The thermostat controls how much heat comes out (set higher = more heat)
In reality, thermostats work by switching the heat on at full power until the desired temperature is reached, then turning off completely. Setting the thermostat at one extreme cannot affect how long it takes to reach the desired temperature. The design gives absolutely no hint as to the correct answer, so people form theories (mental models) to explain what they observe.
Why do people use erroneous mental models? They are running a mental model based on general valve theory, assuming the principle of "more is more": the more you turn or push something, the more it causes the desired effect. This holds for taps and radio controls but not for thermostats, which function based on an on-off switch principle.
Other common examples of erroneous mental models include:
- People pressing an elevator button multiple times, believing it will make it arrive faster
- Bashing keys when the cursor freezes on a computer screen
- Hitting the top of a TV when it starts acting up
💡 Why this matters: Not having appropriate mental models causes people to become very frustrated, often resulting in stereotypical "venting" behavior. If people could develop better mental models, they would be better equipped to carry out tasks efficiently and know what to do when systems act up.
Self-blaming: When users cannot use an everyday thing, they often blame themselves. If they believe others can use the device and that it is not very complex, they conclude difficulties must be their own fault. This creates a conspiracy of silence, maintaining feelings of guilt and helplessness. Interestingly, people normally attribute their own problems to the environment and other people's problems to their personalities—but with everyday objects, they reverse this tendency.
Reason for self-blaming:
- Learned helplessness: People experience failure at a task numerous times, decide the task cannot be done (at least not by them), and stop trying. A few bad experiences can create this feeling, potentially leading to depression.
- Taught helplessness: The design of everyday things seems almost guaranteed to cause this phenomenon. With badly designed objects, faulty mental models, and poor feedback, people feel guilty. The vicious cycle: fail at something → think it's your fault → think you can't do it → don't try → can't do it → self-fulfilling prophecy.
The nature of human thought and explanation: It isn't always easy to tell where blame should be placed. Dramatic accidents have occurred from false assessment of blame. When instruments indicate something is wrong, operators must consider whether the instruments themselves are wrong.
📌 Example — Three Mile Island Nuclear Power Plant: Operators pushed a button to close a valve that had been properly opened. The valve was deficient and didn't close, but a light indicated the valve position was closed. The light monitored only the electrical signal, not the actual valve. Operators saw high temperature in the pipe leading from the valve, indicating fluid was still flowing, but knew the valve had been leaky and assumed the leak wouldn't affect main operation. They were wrong—escaping water added significantly to the nuclear disaster.
📌 Example — Lockheed L-1011: Flying from Miami to Nassau, the low oil pressure light for one engine went on. The crew shut it down and turned back. Eight minutes later, low-pressure lights for the other two engines went on. The crew didn't believe it—simultaneous oil exhaustion in all three engines seemed "one in millions." The second and third engines failed for lack of oil. Three missing O-rings, one missing from each oil plug, allowed all oil to seep out. The National Transportation Safety Board declared the crew's analysis was "logical."
💡 Why this matters: Once we have an explanation—correct or incorrect—for discrepant events, there is no more puzzle. We become complacent. Our explanations are based on analogy with past experience that may not apply in the current situation.
How people do things: To get something done, you need:
- The goal—what is to be achieved
- Action to move yourself or manipulate something
- Check to see that your goal was met
A goal must be transformed into specific statements called intentions. An intention is the specific action taken to get to the goal. Intentions must be made into specific action sequences that control muscles.
📌 Example: Sitting in an armchair reading as light dims. Goal: get more light. Intention: push the switch button on the lamp. Action sequence: how to move the body, stretch to reach the switch, extend the finger to push the button.
Action Cycle: Human action has two aspects—execution and evaluation. Execution involves doing something. Evaluation is the comparison of what happened in the world with what we wanted to happen.
Stages of Execution:
- Start with the goal (state to be achieved)
- Goal translated into an intention to do some action
- Intention translated into a set of internal commands, an action sequence
- Action sequence is executed, performed upon the world
Stages of Evaluation:
- Perception of the world
- Perception is interpreted according to expectations
- Interpretation is compared with respect to both intentions and goals
Seven stages of action: The stages of execution (intentions, action sequence, and execution) are coupled with the stages of evaluation (perception, interpretation, and evaluation), with goals common to both stages.
11.2 Errors
Human capability for interpreting and manipulating information is impressive, but we do make mistakes. When learning to use a computer system, learners are often frightened of making errors because, as well as making them feel stupid, they think it can result in catastrophe. The anticipation of making an error can hinder a user's interaction with a system.
Some errors result from changes in the context of skilled behavior. If a pattern of behavior has become automatic and we change some aspect, the more familiar pattern may break through. Example: intending to stop at the shop on the way home but driving past—the more familiar activity of driving home overrides the less familiar intention.
Other errors result from an incorrect understanding or model of a situation or system—these are mental models. Characteristics: partial, unstable, subject to change, internally inconsistent, often unscientific, based on superstition or incorrect interpretation of evidence.
A classification of errors: Norman categorized errors into two main types:
🔑 Definition — Mistakes: Occur through conscious deliberation. An incorrect action is taken based on an incorrect decision. Example: trying to throw the icon of the hard disk into the wastebasket as a way of removing all existing files from the disk is a mistake. A menu option to erase the disk is the appropriate action.
🔑 Definition — Slips: Are unintentional. They happen by accident, such as making typos by pressing the wrong key or selecting the wrong menu item by overshooting. The most frequent errors are slips, especially in well-learned behavior.
⭐ Key Takeaways
Mental models are incomplete, unstable representations that people form through experience and instruction to predict and understand system behavior. Erroneous mental models, like the "more is more" philosophy applied to thermostats and elevator buttons, are surprisingly common and lead to user frustration and self-blame. The phenomenon of learned helplessness and taught helplessness explains why repeated failure with poorly designed objects creates self-fulfilling prophecies of inability. The seven stages of action (goal, intention, action sequence, execution, perception, interpretation, evaluation) provide a framework for understanding how people interact with systems. Finally, errors are classified into mistakes (conscious incorrect decisions) and slips (unintentional accidents), with slips being the most frequent in well-learned behavior.
🧠 Quick Revision Questions
- What are the key characteristics of mental models according to Norman and Johnson-Laird?
- Why do people set thermostats to maximum when trying to heat a room quickly, and what is the actual correct functioning of a thermostat?
- What is the difference between learned helplessness and taught helplessness in the context of human-computer interaction?
- What are the seven stages of action in the action cycle, and how do execution and evaluation differ?
- What is the difference between a mistake and a slip, and provide an example of each?
📘 Lecture 12 — Design Principles
📖 Overview: This lecture focuses on conceptual models in HCI and key design principles that guide the creation of usable interfaces. It explains how designers can bridge the gap between user mental models and system implementation through represented models, and introduces six fundamental design principles that make interfaces intuitive and easy to use.
🗂️ Topics Covered
The lecture begins by defining conceptual models and their importance in design, then explores the three models in digital systems (user mental model, designer's model, and implementation model). It discusses the Gulfs of Execution and Evaluation, introduces the Seven Stages of Action as a design aid, and concludes with a detailed examination of six core design principles: Visibility, Affordance, Constraints, Mapping, Consistency, and Feedback.
📝 Lecture Summary
Conceptual Model
The lecture opens with a quote from David Liddle: "The most important thing to design is the user's conceptual model." A conceptual model is defined as a description of the proposed system in terms of a set of integrated ideas and concepts about what it should do, behave, and look like, that will be understandable by the users in the manner intended. Developing a conceptual model involves envisioning the proposed product based on user needs and requires iterative testing. A key aspect is deciding what the user will be doing—searching, creating documents, communicating, or recording events. The interaction mode (browsing vs. direct questioning) must be considered separately from the interaction style (menu-based, speech inputs, commands), with the former being at a higher level of abstraction. Once possible ways of interacting are identified, the conceptual model is "fleshed out" by working out interface behavior, interaction style, and "look and feel." Another approach is using interface metaphors, like the desktop, which provide a basic structure using familiar knowledge.
The lecture then introduces three models in the digital world:
- Implementation model — what is actually happening inside the computer
- Designer's represented model — the way the designer chooses to represent a program's functioning to the user
- User's mental model — the user's understanding of how the system works
Donald Norman refers to the represented model simply as the designer's model. In software, a program's represented model can be quite different from the actual processing structure. For example, an operating system can make a network file server look like a local disk even though the physical disk drive may be miles away. The closer the represented model comes to the user's mental model, the easier the program is to use. Offering a represented model that follows the implementation model too closely reduces the user's ability to learn and use the program. We form mental models that are simpler than reality, so creating represented models that are simpler than the implementation model helps users achieve better understanding—like imagining pressing a brake pedal pushes a lever against wheels, when the actual mechanism involves hydraulic cylinders and perforated disks.
🔑 Definition — Conceptual model: A description of the proposed system in terms of a set of integrated ideas and concepts about what it should do, behave, and look like, that will be understandable by the users in the manner intended.
🔑 Definition — Interface metaphor: A basic structure for the conceptual model that is couched in knowledge users are familiar with, such as the desktop metaphor.
🔑 Definition — Designer's represented model: The way the designer chooses to represent a program's functioning to the user, which can differ from the actual processing structure.
The Gulf of Execution and Gulf of Evaluation
There are several gulfs that separate mental states from physical ones, each reflecting the distance between the person's mental representation and the physical components of the environment. These present major problems for users.
The Gulf of Execution — Does the system provide actions that correspond to the person's intentions? The difference between intentions and allowable actions is the gulf of execution. One measure is how well the system allows the person to do intended actions directly, without extra effort.
The Gulf of Evaluation — Does the system provide a physical representation that can be directly perceived and interpreted in terms of the person's intentions and expectations? The gulf of evaluation reflects the effort required to interpret the physical state of the system and determine how well expectations have been met. The gulf is small when the system provides information that is easy to get, easy to interpret, and matches how the person thinks of the system.
The Seven Stages of Action serves as a valuable design aid, providing a basic checklist to ensure the gulfs are bridged. Each stage of action requires its own design strategies and provides its own opportunity for disaster. The questions for each stage boil down to the principles of good design.
🔑 Definition — Gulf of Execution: The difference between the person's intentions and the allowable actions provided by the system.
🔑 Definition — Gulf of Evaluation: The amount of effort the person must exert to interpret the physical state of the system and determine how well expectations and intentions have been met.
12.1 Design Principles
The lecture describes six common design principles: Visibility, Affordance, Constraints, Mapping, Consistency, and Feedback.
Visibility
The more visible functions are, the more likely users will be able to know what to do next. When functions are "out of sight," they become more difficult to find and know how to use. Norman describes car controls as an example—indicator, headlights, horn, and hazard warning lights are clearly visible and positioned to make it easy for the driver to find the appropriate control. The lecture provides a personal example of encountering a problem in word processing software where the "Properties" option was hidden under an arrow at the bottom of the File menu, requiring the user to click the arrow to see it—demonstrating poor visibility.
Affordance
Affordance refers to an attribute of an object that allows people to know how to use it. For example, a mouse button invites pushing by the way it is physically constrained in its plastic shell. At a simple level, to afford means "to give a clue." When affordances are perceptually obvious, it is easy to know how to interact—a door handle affords pulling, a cup handle affords grasping, a mouse button affords pushing.
There are two kinds of affordance:
- Real affordances — Physical objects have real affordances like grasping that are perceptually obvious and do not have to be learned.
- Perceived affordances — Screen-based interfaces are better conceptualized as perceived affordances, which are essentially learned conventions. Physical devices like control consoles benefit from real affordances (pulling, pressing), while virtual interfaces rely on perceived affordances such as icons that afford clicking and scroll bars that afford moving up and down.
💡 Why this matters: Understanding the difference between real and perceived affordances helps designers create interfaces where users intuitively know what to do without instruction.
Constraints
Constraints refer to determining ways of restricting the kind of user interaction that can take place at a given moment. A common design practice is deactivating certain menu options by shading them, restricting users to only permissible actions and reducing the chances of mistakes. Different kinds of graphical representations can also constrain interpretation—for example, flow chart diagrams show which objects are related to which.
Norman classified constraints into three categories:
-
Physical constraints — The way physical objects restrict movement. For example, a disk can be inserted into a disk drive in only one way due to its shape and size. Keys on a pad can usually be pressed in only one way.
-
Logical constraints — Rely on people's understanding of how the world works and common-sense reasoning about actions and consequences. For example, picking up a physical marble and placing it elsewhere would be expected to trigger something. Disabling menu options when inappropriate provides logical constraining.
-
Cultural constraints — Rely on learned conventions, like using red for warning, certain signals for danger, or the smiley face for happy emotions. Most cultural constraints are arbitrary—their relationship with what is being represented is abstract and could have evolved differently (e.g., yellow instead of red for warning). Once learned and accepted by a cultural group, they become universally accepted conventions. Two universally accepted interface conventions are windowing for displaying information and icons on the desktop.
🔑 Definition — Constraints: Ways of restricting the kind of user interaction that can take place at a given moment.
Mapping
Mapping refers to the relationship between controls and their effects in the world. Nearly all artifacts need some kind of mapping between controls and effects. An example of good mapping is the up and down arrows on a computer keyboard representing up and down cursor movement. The mapping of relative position is also important—consider music players where play is in the middle, rewind on the left, and fast-forward on the right, mapping directly onto the directionality of actions. The lecture illustrates this with two figures: Figure (a) shows poor mapping (which would be difficult to use) and Figure (b) shows good mapping following the conventional sequence.
🔑 Definition — Mapping: The relationship between controls and their effects in the world.
Consistency
Consistency refers to designing interfaces to have similar operations and use similar elements for achieving similar tasks. A consistent interface follows rules, such as using the same operation (left mouse button click) to select all objects. Inconsistent interfaces allow exceptions to a rule—for example, where certain graphical objects are highlighted using the right mouse button while all others use the left button. This makes it difficult for users to remember and prone to mistakes.
Benefits of consistent interfaces include being easier to learn and use—users learn only a single mode of operation applicable to all objects. This works well for simple interfaces with limited operations (like a mini CD player). However, consistency is more problematic for complex interfaces with hundreds of operations, where creating categories of commands mapped into subsets of operations is more effective (like categorizing word-processing operations into different menus).
Another problem is determining what to make consistent with what else. There are many choices:
- External consistency — Designing to be consistent with how people do things in the outside world
- Internal consistency — Designing to be consistent with the existing system
The problem facing designers is knowing which approach to use, as there are many different ways of doing things both physically and electronically.
🔑 Definition — Consistency: Designing interfaces to have similar operations and use similar elements for achieving similar tasks.
Feedback
Feedback is related to visibility and involves sending back information about what action has been done and what has been accomplished, allowing the person to continue with the activity. Without feedback, actions would have no effect—imagine playing a guitar, slicing bread, or writing with a pen if none of the actions produced any effect for several seconds. Various kinds of feedback are available for interaction design: audio, tactile, verbal, visual, and combinations of these. Deciding which combinations are appropriate for different activities is central. Using feedback in the right way also provides the necessary visibility for user interaction.
🔑 Definition — Feedback: Sending back information about what action has been done and what has been accomplished, allowing the person to continue with the activity.
⭐ Key Takeaways
The most critical concept from this lecture is that a well-designed conceptual model bridges the gap between the user's mental model and the system's implementation through an appropriate represented model. The six design principles—Visibility, Affordance, Constraints, Mapping, Consistency, and Feedback—serve as practical guidelines for creating interfaces that are intuitive and easy to use. Understanding the Gulfs of Execution and Evaluation helps designers identify where users struggle, while the Seven Stages of Action provides a checklist for bridging these gulfs. Students must remember the three types of constraints (physical, logical, cultural) and the difference between real and perceived affordances, as these are fundamental to designing interfaces that prevent errors and guide user behavior naturally.
🧠 Quick Revision Questions
- What is the difference between the user's mental model, the designer's represented model, and the implementation model?
- How do the Gulf of Execution and the Gulf of Evaluation differ, and what does it mean when each gulf is "small"?
- What are the three categories of constraints according to Norman, and give one example of each from user interfaces?
- Why can consistency be problematic when designing complex interfaces with hundreds of operations?
- What is the difference between real affordances and perceived affordances, and which type is more relevant for screen-based interfaces?
📘 Lecture 13 — The Computer
📖 Overview: This lecture shifts focus from human aspects to computer aspects of Human Computer Interaction. It provides a comprehensive overview of input and output devices, examining their design, usability, and suitability for different tasks and users. Understanding these devices is crucial for designing systems that are safe, effective, efficient, and enjoyable to use.
🗂️ Topics Covered
This lecture begins by defining input devices and their selection criteria, emphasizing the importance of matching devices to user physiology, task, and environment. It then explores various text entry devices such as keyboards (QWERTY, alphabetic, Dvorak, chord), phone pads with T9, handwriting, and speech recognition. The lecture continues with a detailed discussion of positioning and pointing devices including the mouse, touch pad, trackball, joystick, touch screen, stylus, and eyegaze. It concludes with an overview of display devices (CRT, LCD, digital paper), sound output, tactile feedback, physical controls, and environmental sensing.
📝 Lecture Summary
13.1 Input devices
Input is concerned with recording and entering data into a computer system and issuing instructions. Input devices can be defined as a device that, together with appropriate software, transforms information from the user into data that a computer application can process. The key aim in selecting an input device is to help users carry out their work safely, effectively, efficiently, and enjoyably. The most appropriate device will match the physiology and psychological characteristics of users, be appropriate for the tasks to be performed (e.g., continuous movement for drawing, discrete movement for selecting from a list), and be suitable for the intended work and environment.
Frequently, no single optimal device can be identified, and trade-offs must be made. Many systems use two or more input devices together, which must be complementary and well coordinated. There must also be adequate and appropriate system feedback to guide users. Feedback can be visual (text appearing, cursor moving), auditory (an alarm, keys clicking), or tactile (using a joystick), often in combination.
🔑 Definition — System Feedback: Responses from the system to guide, reassure, inform, and correct user’s errors. It can be visual, auditory, or tactile.
13.2 Text entry devices
Keyboard The most common method of entering information is through a keyboard. Broadly defined, a keyboard is a group of on-off push buttons. It is a discrete entry device, sensing one of two or more discrete positions, as opposed to continuous entry devices which sense in a continuous range. The physical design of individual keys and their grouping arrangements are important. For example, keys that are too small cause difficulty in hitting them accurately. Membrane keyboards are sealed and can withstand grease, dirt, and liquids, making them suitable for environments like production floors, but they lack tactile feedback. Alterations in key arrangement affect speed and accuracy. Research shows trained typists look ahead and process text in chunks of two to three words for alphabetic text and three to four characters for numerical material.
🔑 Definition — Discrete Entry Device: An input device that senses one of two or more discrete positions (e.g., keyboard keys, touch-sensitive switches). 🔑 Definition — Continuous Entry Device: An input device that senses in a continuous range (e.g., pens with digitizing tablets, moving joysticks).
QWERTY keyboard
This standard alphanumeric layout, named after the first letters in the uppermost row, became a commercial success in 1874. The arrangement was chosen to reduce keys jamming in manual typewriters, not for optimal typing. For example, the letters s, t, and h are far apart even though they are frequently used together.
Alphabetic keyboard This layout arranges letters alphabetically. Despite expectations, it is not faster for properly trained typists and makes little difference to the speed for novice or occasional users. It is sometimes used in pocket electronic organizers to dissuade touch-typing on a very small keyboard.
Dvorak Keyboard Patented in 1932, the Dvorak board was designed based on the frequency of letter usage and patterns in English. All vowels and the most frequent consonants are on the home row, so approximately 70% of common words are typed on this row. It promotes faster operation by tapping with alternate hands. Dvorak claimed this arrangement reduces between-row movement by 90%. Despite its benefits, it has never been commercially successful due to the cost of replacing existing keyboards and retraining users.
📌 Example: On a Dvorak keyboard, the vowels a, o, e, and u and the most frequent consonants are typed on the home row, reducing finger travel.
Chord keyboards On chord keyboards, several keys must be pressed at once to enter a single character, similar to playing a flute. Few keys are required, so they can be very small and operated with one hand. They are useful where space is limited or one hand is occupied. Training is required to learn the finger combinations. They are used for mail sorting and recording court transcripts.
Phone pad and T9 entry
Mobile phones use a keypad with digits 0-9, not a full keyboard. For text input, numeric keys are pressed multiple times. For example, the 3 key has def; pressing it once gives d, twice gives e, three times gives f. The T9 algorithm uses a large dictionary to disambiguate words by typing each letter only once. For example, 3926753 becomes example as it is the only real word that matches.
📌 Example: To type the word "example" using T9, you press 3926753. The phone's dictionary recognizes this sequence as "example" because alternative letter combinations (like ewbosld) are not real words.
Handwriting recognition Handwriting is a familiar activity, making it an attractive text entry method. However, current technology is often inaccurate and individual differences in handwriting are enormous. The most significant information is in the stroke information—the way the letter is drawn—so devices must capture this. Online recognition is easier than reading text on paper. A key limitation is speed, as it is difficult to write more than 25 words per minute, which is half the speed of a decent typist. Pen-based systems are popular in mobile computing for taking notes and sketching, as they can be small and accurate, unlike small keyboards.
Speech recognition Speech input offers several advantages: it is a natural form of communication, does not require hands, and aids disabled users. However, it suffers from several problems: it is limited to specialized tasks, can be inaccurate in distinguishing similar-sounding words, is subject to background noise, and natural language is hard for computers to interpret. Systems range from isolated word recognition (requiring pauses between words) to continuous speech recognition (allowing faster, more natural input). Speaker-dependent systems require training, while speaker-independent systems accommodate a large range of speaking characteristics but are less reliable.
🔑 Definition — Continuous Speech Recognition System: A system capable of recognizing words spoken in a natural, flowing manner without pauses between words, allowing faster data entry.
13.3 Positioning, Pointing and Drawing
Pointing devices are input devices used to specify a point or path in space. Examples include the mouse, touch pad, trackball, joystick, touch screen, and eye gaze.
Mouse The mouse is a small, palm-sized box with a weighted ball. As the box is moved, the ball rotates inside, and this rotation is detected by rollers. It is an indirect input device because a transformation maps horizontal desktop motion to vertical screen alignment. Left-right motion is directly mapped, while up-down is achieved by moving the mouse towards or away from the user.
🔑 Definition — Indirect Input Device: An input device for which a transformation is required to map its movement to the screen's output (e.g., mouse to screen cursor).
Foot mouse A foot-operated device, akin to an isometric joystick, that allows hands to be dedicated to the keyboard. It is a rare device.
Touch pad Touchpads are touch-sensitive tablets, usually 2-3 inches square, operated by stroking a finger over their surface. Because they are small, they may require several strokes to move the cursor across the screen. This can be improved by acceleration settings in the software, where the pad-to-screen distance ratio varies with the speed of finger movement.
📌 Example: If you move your finger slowly on a touchpad, it maps to small screen movements. If you move your finger quickly, the same distance on the pad moves the cursor a long distance on the screen.
Trackball and thumbwheel A trackball is an upside-down mouse with a weighted ball in a static housing. It is compact, requires no additional space, but is hard for drawing. Thumbwheels have two orthogonal dials to control the cursor position; they are cheap but slow and difficult to use for diagonal movement. Single thumbwheels are often included on mice for scrolling documents.
Joystick and track point The joystick is an indirect input device consisting of a stick in a box. An absolute joystick has its stick position correspond to the cursor position. An isometric joystick uses pressure on the stick to control cursor velocity, and the stick returns to center when released. A track point is a smaller version used on laptops, often a rubber nipple in the center of the keyboard.
Touch screens Touch screens allow input by touching the screen, making it a bi-directional device (input and output). Advantages are ease of learning, durability, and direct interaction. Disadvantages include greasy marks on the screen, inaccuracy (making small selections and drawing difficult), and arm fatigue from reaching to a vertical screen.
Stylus and light pen A stylus is a pen-like stick for pointing and drawing, popular in PDAs. A light pen is an older technology connected to the screen by a cable. It detects a burst of light from the screen phosphor during the display scan, allowing it to address individual pixels, making it more accurate than a touch screen.
Eyegaze Eyegaze systems control the computer by tracking eye movement. A low-power laser is shone into the eye, and the reflection changes as the eye angle alters. The system determines the direction of the gaze. It is fast and accurate but can be expensive. It is fine for selection but not for drawing because the eye does not move in smooth lines, and it can be difficult to distinguish a deliberate gaze from an accidental glance.
Cursor keys
Four keys on the keyboard (up, down, left, right) are used to control the cursor. There is no standardized layout, but the most common is the inverted T. They were more heavily used in character-based systems.
13.4 Display devices
Cathode ray tube The CRT is a television-like screen. An electron gun emits a stream of electrons, which are focused and directed by magnetic fields to hit a phosphor-coated screen, causing it to glow. The beam is scanned from left to right, top to bottom. Color is achieved using three electron guns for red, green, and blue phosphors. CRTs are cheap with fast response times and high color capability but are bulky.
Liquid Crystal Display LCDs are light, flat plastic screens using liquid crystal technology. They are smaller, lighter, and consume far less power than CRTs. They are matrix addressable, meaning individual pixels can be accessed without scanning. They have no radiation problems and are common in notebook and portable computers.
🔑 Definition — Matrix Addressable: A display technology where individual pixels can be accessed directly without the need for scanning the entire screen.
Digital paper A new form of display that is thin, flexible, and can be written to electronically. It retains its contents even when removed from an electrical supply.
Sound output Auditory signals are used in conjunction with screen displays for system output and feedback. Sounds like beeps, clicks, and tones convey information. For example, keyboards can emit a click when a key is pressed, speeding up performance.
13.5 Touch, feel and smell
The sense of touch and feel is used for tactile feedback. Technology for simulating textures is just beginning to become available.
13.6 Physical controls
Dedicated control panels are designed for a single device and use, unlike generic desktop computer controls. For example, a microwave uses a flat plastic control panel because it is used in a kitchen with greasy hands. The smooth surface has no gaps for food to accumulate and is easy to clean.
📌 Example: A microwave has a smooth, flat control panel with no gaps, making it easy to clean, while a washing machine has large buttons that act as both controls and displays, as hands are less greasy.
13.7 Environment and bio sensing
Many sensors exist in our environment, such as for automatic doors and energy-saving lights. The vision of ubiquitous computing suggests our world will be filled with such devices that monitor our behavior.
⭐ Key Takeaways
The choice of input and output devices is a critical design decision in HCI, requiring a trade-off between user characteristics, task requirements, and environmental constraints. For text input, the QWERTY keyboard remains dominant despite its suboptimal design, while alternatives like Dvorak offer ergonomic benefits but face adoption barriers. For pointing, the mouse is a standard, but other devices like touch screens and trackpads are better suited for specific contexts (e.g., public kiosks, laptops). System feedback—visual, auditory, and tactile—is essential for guiding users and preventing errors. Finally, emerging technologies like speech recognition and eyegaze offer new interaction paradigms, particularly for users with disabilities, though they still face significant limitations in accuracy and naturalness.
🧠 Quick Revision Questions
- Explain the three key factors to consider when selecting an input device for an interactive system.
- Compare and contrast the QWERTY, Dvorak, and alphabetic keyboard layouts, focusing on their design principles and usability.
- Describe the difference between an absolute joystick and an isometric joystick.
- What are the primary advantages and disadvantages of using a touch screen as an input device?
- Explain the difference between speaker-dependent and speaker-independent speech recognition systems and give an example of a problem each might face.
📘 Lecture 14 — Interaction
📖 Overview: This lecture introduces the concept of interaction in Human Computer Interaction, explaining how humans and computers communicate. It covers key terminology, theoretical models of interaction, and various interaction styles, providing a foundational understanding of how users engage with computer systems.
🗂️ Topics Covered
The lecture defines core terms of interaction including domain, task, and goal, then presents Donald Norman’s model of interaction with the seven stages of action and the gulfs of execution and evaluation. It introduces an interaction framework breaking the system into four components (System, User, Input, Output) and discusses how frameworks relate to HCI, focusing on ergonomics, physical aspects of interfaces, industrial interfaces, glass interfaces, indirect manipulation, and the physical environment. Finally, it describes various interaction styles including command line interface, menus, natural language, question/answer and query dialog, form-fills and spreadsheets, WIMP, point and click, and three-dimensional interfaces.
📝 Lecture Summary
14.1 The terms of Interaction
Domain defines an area of expertise and knowledge in some real-world activity. Examples include graphic design, authoring, and process control in a factory. A domain consists of concepts that highlight its important aspects. In a graphic design domain, important concepts are geometric shapes, a drawing surface, and a drawing utensil.
Tasks are the operations to manipulate the concepts of a domain. For example, constructing a specific geometric shape with particular attributes on the drawing surface is a task within the graphic design domain.
Goal is the desired output from a performed task. A related goal would be to produce a solid red triangle centered on the canvas. So, the goal is the ultimate result you want to achieve after performing specific tasks.
14.2 Donald Norman’s Model
Donald Norman’s Model of interaction describes how a user chooses a goal, formulates a plan of action, which is then executed at the computer interface. When the plan has been executed, the user observes the computer interface to evaluate the result and determine further actions.
The two major parts, execution and evaluation, of the interactive cycle are subdivided into seven stages. For example: you decide you need more light (goal), form an intention to switch on the desk lamp, specify actions to reach over and press the switch, execute the action, perceive the result (light on or off), interpret this based on knowledge of the world, and evaluate against the original goal.
🔑 Definition — Gulf of execution: the difference between the user’s formulation of the actions to reach the goal and the actions allowed by the system. If the actions allowed correspond to those intended by the user, the interaction will be effective.
🔑 Definition — Gulf of evaluation: the distance between the physical presentation of the system state and the expectation of the user. If the user can readily evaluate the presentation in terms of their goal, the gulf of evaluation is small.
💡 Why this matters: These two gulfs represent common interface problems. Good interface design aims to minimize both gulfs to make interaction more intuitive and effective.
14.3 The interaction framework
The interaction framework breaks the system into four main components: the System, the User, the Input, and the Output. Each component has its own language. The system’s language is the core language (describing computational attributes of the domain relevant to system state), and the user’s language is the task language (describing psychological attributes of the domain relevant to user state). Input and output together form the interface.
There are four steps in the interactive cycle, each corresponding to a translation between components: articulation (user formulates task and articulates it in input language), performance (input language is translated into core language as operations), presentation (system state is rendered as output features), and observation (user observes output and assesses results relative to the original goal).
14.4 Frameworks and HCI
The ACM SIGCHI Curriculum Development Group presents a framework placing different areas relating to HCI. Ergonomics addresses issues on the user side of the interface, covering input and output as well as the user’s immediate context. Dialog design and interface styles are placed along the input branch, while presentation and screen design relate to the output branch.
Ergonomics (or human factors) is the study of the physical characteristics of interaction: how controls are designed, the physical environment, and the layout and physical qualities of the screen. Physical aspects of interface include:
- Arrangement of controls and displays: Users should group sets of controls and parts of the display logically for rapid access, especially in safety-critical applications.
- Industrial Interface: Industrial interfaces may require rapid assimilation of multiple numeric displays varying in response to the environment, raising additional design issues.
- Glass interfaces vs. dials and knobs: A glass interface (computer screen) can be cheaper and more flexible for complex systems, allowing information to be shown in multiple forms.
- Indirect manipulation: Industrial interfaces are intermediaries between operator and the real world, requiring feedback at two levels: immediate feedback that actions were received, and monitoring of equipment effects.
The physical environment of the interaction considers where the system will be used and by whom. Health issues include physical position (comfortable reach and support), temperature (extremes affect performance), lighting (adequate without glare), noise (maintain comfortable levels), and time (control usage duration).
Use of color: The human visual system has limitations regarding color, including the number of distinguishable colors and relatively low blue acuity. Colors should be distinct, not affected by contrast changes, and correspond to common conventions and user expectations, though color conventions are culturally determined.
14.5 Interaction styles
Interaction is communication between computer and human. For successful enjoyable communication, interface style has its own importance. Common interface styles include:
-
Command line interface: The first interactive dialog style, providing direct access to system functionality using function keys, single characters, abbreviations, or whole-word commands. They are powerful, flexible with options/parameters, and useful for repetitive tasks.
-
Menu: Options available to the user are displayed on screen and selected using mouse or keys. Since options are visible, they rely on recognition rather than recall. Menus are often hierarchically ordered; grouping and naming provides cues for finding options.
-
Natural Language: Attractive but difficult because the ambiguity of natural language makes it very difficult for a machine to understand.
-
Question/answer and query dialog: Question/answer leads users through interaction step by step — easy to learn but limited in functionality. Query languages construct queries to retrieve information from databases using natural-language-style phrases but require specific syntax and knowledge of database structure.
-
Form-fills and spreadsheets: Form-filling interfaces present a display resembling a paper form with slots to fill in, allowing easy movement and correction facilities. Spreadsheets comprise a grid of cells, each containing a value or formula. Modern example: MS Excel; historical examples: VISICALC and Lotus 123.
-
The WIMP Interfaces: Stands for Windows, Icons, Menus, and Pointers. This is the default interface style for most interactive computer systems today, especially in PC and desktop workstation environments.
-
Point and Click interface: In multimedia systems and web browsers, virtually all actions take only a single click. Pointing at a city on a map and clicking opens tourist information; pointing at a word shows a definition.
-
Three-dimensional interfaces: Increasing use of 3D effects in user interfaces. The most obvious example is virtual reality (VR), but simpler techniques include giving ordinary WIMP elements a 3D appearance using shading.
⭐ Key Takeaways
The lecture establishes that interaction is the core of HCI — the communication between human and computer users. Norman’s model with its seven stages of action and the concepts of gulf of execution and gulf of evaluation are critical for understanding how users interact with systems and for identifying interface problems. The interaction framework with its four components (System, User, Input, Output) and four translations (articulation, performance, presentation, observation) provides a systematic way to analyze interaction. Ergonomics and physical considerations including health issues, environmental factors, and color use directly impact user performance and safety. Finally, the various interaction styles (from command line to WIMP to 3D interfaces) each have distinct advantages and limitations that make them appropriate for different contexts and user types.
🧠 Quick Revision Questions
- What are the four main components in the interaction framework, and what are the four translations involved in the interactive cycle?
- Explain the difference between the gulf of execution and the gulf of evaluation in Norman’s model. Provide an example of each.
- What are the seven stages of action in Norman’s model of interaction? Use the desk lamp example to illustrate them.
- List and briefly describe at least five different interaction styles covered in this lecture.
- What physical environment and health issues should be considered in ergonomic design of interfaces?
📘 Lecture 15 — Interaction Paradigms
📖 Overview: This lecture provides a detailed examination of WIMP interfaces—Windows, Icons, Pointers, and Menus—along with related widgets like toolbars, buttons, and dialog boxes. It then explores the evolution of interaction paradigms, from time-sharing and personal computing to modern concepts like ubiquitous computing and the World Wide Web, showing how technological advances have reshaped human-computer interaction.
🗂️ Topics Covered
This lecture covers two main sections: first, a detailed breakdown of the WIMP interface components including windows, icons, pointers, menus, keyboard accelerators, buttons, radio buttons, check boxes, toolbars, palettes, and dialog boxes. The second section introduces interaction paradigms, discussing time-sharing, video display units, programming toolkits, personal computing, WIMP interface history, the metaphor, direct manipulation, language versus action paradigms, hypertext, multi-modality, computer-supported cooperative work, the World Wide Web, ubiquitous computing, and sensor-based interaction.
📝 Lecture Summary
15.1 The WIMP Interfaces
The lecture begins by revisiting and expanding on the four key features of the WIMP interface: windows, icons, pointers, and menus. Beyond these, many additional interaction objects and techniques are commonly used, including toolbars, menus, buttons, palettes, and dialog boxes. Collectively, these elements are called widgets, and they form the toolkit for interaction between user and system.
Windows Windows are areas of the screen that behave as if they were independent terminals. They can contain text or graphics, be moved or resized, and allow multiple tasks to be visible simultaneously. When one window overlaps another, the back window is partially obscured and refreshed when exposed again. Overlapping windows can obscure vital information, so windows may also be tiled (adjoining without overlapping) or placed in a cascading fashion (each new window offset slightly down and to the left). Scrollbars allow users to move window contents up/down or side to side, making the window behave like a real window onto a larger world. A title bar at the top identifies the window, and corner boxes aid resizing, closing, or maximizing. Some systems also allow windows within windows, such as in Microsoft Office where each application has its own window and each document has a sub-window.
Icons Windows can be closed and lost forever, or they can be shrunk to a much-reduced representation called an icon. By allowing icons, many windows can be available on the screen simultaneously, ready to be expanded by clicking. Iconifying a window suspends that dialog temporarily, saving screen space and serving as a reminder that the dialog can be resumed. Icons can also represent other system aspects (e.g., a wastebasket, disks, programs). Icons can be realistic, stylized, or arbitrary symbols, though arbitrary symbols can be difficult for users to interpret.
Pointers The pointer is essential because WIMP interaction relies heavily on pointing and selecting. The mouse provides input, though joysticks and trackballs are alternatives. The user sees a cursor on the screen controlled by the input device. Different cursor shapes distinguish modes (e.g., arrow for normal, cross-hairs for drawing). Cursors also indicate system activity (e.g., an hourglass cursor when the system is busy). Pointer cursors have a hot-spot, the exact location to which they point.
Menus A menu presents a choice of operations or services available at a given time. Since recall is inferior to recognition, menus provide visual cues in an ordered list. The pointing device highlights items as the pointer moves over them, and selection usually requires pressing a mouse button or keyboard key. Cascading menus allow refinement by opening another menu adjacent to the selected item. Main menus can be permanently visible as a menu bar (often at screen/window top) or hidden as pop-up menus that appear upon request, often context-sensitive. Pull-down menus require pressing a button to drag down, while fall-down menus appear automatically when the pointer enters the title bar. Pie menus arrange options in a circle with the pointer in the center, making selection distance equal for all items, though they take more screen space. The major problems with menus are deciding what items to include and how to group them; menu labels should reflect function, items grouped by function, groupings consistent across applications, and items ordered by importance and frequency of use.
Keyboard accelerators Keyboard accelerators are key combinations that have the same effect as selecting a menu item, allowing expert users to avoid moving off the keyboard. Accelerators are often displayed alongside menu items so frequent use makes them familiar.
Buttons Buttons are individual, isolated regions within a display that can be selected to invoke specific operations. They are designed to resemble push buttons on a control panel. "Pushing" a button invokes a command, usually indicated by a textual label or icon.
Radio Buttons Radio buttons are toggle buttons grouped together to allow selection of one feature from a set of mutually exclusive options (e.g., font size in points).
Check boxes Check boxes are collections of toggle buttons used when a set of options is not mutually exclusive (e.g., bold, italic, underline), indicating the on/off status of each option.
Toolbars Toolbars are collections of small buttons with icons, placed at the top or side of a window, offering commonly used functions. They are similar to menu bars but allow more functions to be displayed simultaneously because icons are smaller than text. Users can often customize the toolbar's content or choose from predefined sets.
Palettes Palettes are a mechanism for making the set of possible modes and the active mode visible to the user. A mode changes the interpretation of user actions (e.g., keystrokes). A palette is usually a collection of icons reminiscent of each mode's purpose, like an artist's palette for paint colors. Users may be able to "tear off" menus or drag toolbars to create palettes placed anywhere on the screen.
Dialog boxes Dialog boxes are information windows used to bring the user's attention to important information (e.g., an error, warning) or to invoke a sub-dialog for a specific task (e.g., saving a file). They factor out auxiliary task threads from the main task dialog.
15.2 Interaction Paradigms
An interaction paradigm is a particular philosophy or way of thinking about interaction design, intended to orient designers to the kinds of questions they need to ask. For many years, the prevailing paradigm was the desktop, with a single user using a GUI/WIMP interface. Recent trends move beyond the desktop with wireless, mobile, and handheld technologies.
Time sharing In the 1940s-1950s, advances were in hardware. By the 1960s, J.C.R. Licklider at ARPA financed research to channel growing computing power. A major contribution was time-sharing, where a single computer supported multiple users. Previously, batch sessions required submitting jobs on punched cards. Time-sharing made programming interactive, creating a subculture of "hackers."
Video display units In the 1950s, researchers experimented with video display units (VDUs). In 1962, Ivan Sutherland's Sketchpad program demonstrated two key ideas: computers could be used for more than data processing, extending the user's ability to abstract and visualize; and one creative mind could make a significant contribution to computing history.
Programming toolkits Douglas Engelbart aimed to use computers to complement human problem-solving. His idea of teaching humans through computers contrasted with the prevailing view that computers were complex and only for the intellectually privileged.
Personal computing Programming toolkits increased productivity for skilled users. The 1970s saw computing aimed at the masses. LOGO was a graphics programming language for children, using a computer-controlled turtle. Alan Kay envisioned small, powerful, single-user machines—personal computers. He worked at Xerox PARC on Smalltalk, a visually based programming environment.
Window systems and the WIMP interface Personal computing focused on single users, but humans think about multiple things at once. A system that forces linear progression does not match human working patterns. The WIMP interface (windows, icons, menus, pointers) became commonplace, first appearing commercially in April 1981 with Xerox's 8010 Star Information System.
The metaphor Metaphor is used to teach new concepts in terms of familiar ones. The Xerox Alto and Star used the desktop metaphor for file manipulation tasks. The spreadsheet is another successful metaphor. However, metaphors can have dangers after the initial honeymoon period (e.g., the typewriter metaphor for word processors, where space is a character, not a passive movement). Metaphors also carry cultural bias that may not apply across national boundaries.
Direct Manipulation Direct manipulation, a term coined by Ben Shneiderman in 1982, describes graphics-based interactive systems. Its features include:
- Visibility of objects of interest
- Incremental action with rapid feedback
- Reversibility of all actions
- Syntactic correctness of all actions
- Replacement of complex command languages with direct manipulation of visible objects
The first commercial success demonstrating direct manipulation for the public was the Apple Macintosh in 1984. The direct manipulation interface for the desktop metaphor makes documents and folders visible as icons; moving a file is mirrored by dragging its icon.
Language versus action Direct manipulation makes some tasks easier but others more difficult. The action paradigm (direct manipulation) involves performing actions at the interface. The language paradigm involves giving instructions via indirect language, with the interface acting as an interlocutor or agent between user and system. The action paradigm is easier for simple tasks; the language paradigm allows describing generic procedures for repeatable tasks. Programming by example combines both, where the user performs actions and the system records them as a generic script.
Hypertext In 1945, Vannevar Bush described the memex, an information storage/retrieval apparatus mimicking human ability to create random associative links. In the 1960s, Nelson coined hypertext to reflect non-linear text structure, inspired by the memex. Traditional text is linear; hypertext allows random, associative browsing.
Multi-modality A multi-modal interactive system relies on multiple human communication channels. Each different channel is a modality. Traditional systems use keyboard and mouse (visual and haptic channels). Genuine multi-modal systems use simultaneous use of multiple channels for both input and output, mimicking natural human information processing.
Computer-supported cooperative work (CSCW) The 1960s saw the first computer networks. As networks became widespread, individuals reconnected their workstations, leading to computer-supported cooperative work (CSCW) . These systems allow interaction between humans via the computer, and designers must consider the society within which users operate. A fine example is electronic mail (email) .
The World Wide Web The World Wide Web (WWW or web) is built on top of the Internet, offering an easy-to-use, graphical interface to information. The Internet is a collection of computers linked by data connections using common transmission protocols. The web adds its own network protocol, a standard markup notation (HTML), and a global naming scheme. Web pages can contain text, images, movies, sound, and most importantly, hypertext links to other pages.
Ubiquitous computing In the late 1980s, Mark Weiser at Xerox PARC initiated a program to move HCI away from the desktop into everyday life. Ubiquitous computing (or pervasive computing) aims to create a computing infrastructure that permeates our physical environment so we no longer notice the computer. The electric motor is a good analogy: it was once large and noticeable, but now it is ubiquitous and invisible.
Sensor-based and context-aware interaction Weiser's dream was "computers anymore." There are increasing technologies that embed computation unobtrusively into daily life, from mobile devices to more pervasive environments.
⭐ Key Takeaways
Students must remember the four core components of the WIMP interface—windows, icons, pointers, and menus—and understand how each functions and their associated issues (e.g., window overlap, icon recognition, menu grouping). The concept of interaction paradigms is critical: these are philosophies that guide design, and they have evolved from time-sharing to personal computing, direct manipulation, and beyond. The metaphor is a powerful but potentially dangerous tool for teaching new concepts. Direct manipulation features (visibility, rapid feedback, reversibility, syntactic correctness, and direct action) contrast with language-based interaction paradigms. Finally, modern paradigms like hypertext, CSCW, the World Wide Web, and ubiquitous computing represent fundamental shifts in how we think about and design human-computer interaction.
🧠 Quick Revision Questions
- What are the four key features of the WIMP interface, and what is the primary purpose of each?
- Explain the difference between pull-down menus, fall-down menus, and pop-up menus.
- What are the five features of a direct manipulation interface as defined by Ben Shneiderman?
- What is the "metaphor" in HCI, and what are its potential dangers?
- What is the core idea behind ubiquitous computing, and how does the electric motor analogy help explain it?
📘 Lecture 16 — HCI Process and Models
📖 Overview: This lecture examines why many software products fail despite technical soundness, emphasizing that poor user experience and usability are the root causes. It introduces the need for new software development models that prioritize user-centered design, and explains the importance of integrating interaction design into the development process to create successful, human-centered products.
🗂️ Topics Covered
The lecture begins by revealing shocking statistics about user failures in software, then explains the true cost of bad user experience versus the rewards of good design. It defines software quality with usability as a key characteristic, distinguishing strategic from tactical aspects. It identifies three core problems in the industry: ignorance about users, conflict of interest between programmers and designers, and the lack of a reliable design process. Finally, it traces the evolution of software development from programmers doing everything to the need for design to precede programming, concluding with the definition and three dimensions of interaction design (form, meaning, behavior).
📝 Lecture Summary
Learning Goals
The aim of this lecture is to introduce you to the study of Human Computer Interaction, so that after studying this you will be able to: • Understand the need of new software development models • Understand the importance of user experience and usability in design
The Cost of Bad User Experience
It has been said, “to err is human; to really screw up, you need a computer.” Inefficient mechanical systems can waste cents, but bad information processes can lose your entire company. Our digital tools are extremely hard to learn, use, and understand, and they often cause us to fall short of our goals, wasting money, time, and opportunity. The irony is that better products don’t take longer or cost more to build; they are difficult only because our process for making them is old-fashioned. Only long-standing traditions rooted in misconceptions keep us from having better products.
Consider a scenario: a website is aesthetically beautiful, technically flawless, and has wonderful animated content. But if a user cannot find desired information or products, it is useless from a business point of view.
📌 Statistics:
- Users can only find information 42% of the time – Jared Spool
- 62% of web shoppers give up looking for the item they want to buy online – Zona Research
- 50% of potential sales from a site are lost because people cannot find the item – Forrester Research
- 40% of users who do not return to a site do so because their first visit resulted in a negative experience – Forrester Research
- 80% of software lifecycle costs occur after release; of that, 80% is due to unmet user requirements, only 20% due to bugs – IEEE Software
- 63% of software projects exceed cost estimates; top reasons include frequent requests for changes from users, overlooked tasks, users' lack of understanding of requirements, and insufficient user-analyst communication – Communications of the ACM
- BOO.com, a $204m startup fails – BBC News
- Poor commercial web sites will kill 80% of Fortune 500 companies within a decade – Jakob Nielsen
All these facts reveal that products with bad user experience deserve to die!
The Reward of Good User Experience
The real importance is of good interaction design. If we design a system with good usability and good user experience, the result will be satisfaction and happiness. People will buy your product and recommend it to others, creating a chain reaction. The premise is simple: if achieving the user’s goal is the basis of our design process, the user will be satisfied and happy, and will gladly pay us money.
🔑 Why this matters: Most digital products emerge from the development process like a monster from a bubbling tank. Developers create technological solutions without planning with users in mind, failing because they have not imbued their creations with humanity.
Software Quality and Usability
We are mainly concerned with software quality, defined as: The extent to which a software product exhibits these characteristics: • Functionality • Reliability • Usability • Efficiency • Maintainability • Portability
Usability can be understood in two aspects:
- Strategic: guides us to think about user interface idioms – the way in which the user and the idiom interact.
- Tactical: gives us hints and tips about using and creating user interface idioms, like dialog boxes and pushbuttons.
Integrating the strategic and tactical approaches is key to designing effective user interaction. There is no objectively good dialog box; quality depends on the situation, the user, and their background and goals. Merely applying tactical dictums doesn’t make the end result better.
Three Reasons Why Products Fail
Ignorance about Users The digital technology industry doesn’t have a good understanding of what it takes to make users happy. Most technology products get built without much understanding of the users. Knowing market segments, income, or jobs does not tell us how users will actually use the product or why they might choose it over competitors.
Conflict of Interest There is an important conflict of interest: the people who build the products—programmers—are usually also the people who design them. Programmers are often required to choose between ease of coding and ease of use. Because performance is judged by coding efficiency and tight deadlines, most software takes the easy-coding direction. Just as we would never permit a prosecutor to also adjudicate a legal case, we should ensure that those designing a product are not the same people building it. It simply isn’t possible for a programmer to advocate for the user, the business, and the technology at the same time.
Lack of a Reliable Process The industry has no reliable, complete process for producing successful products. Engineering follows rigorous methods for technical quality; marketing follows methods for commercial viability. What’s left out is a repeatable, analytical process for transforming an understanding of users into products that meet their needs and excite their imaginations. Most software has never undergone a rigorous design process from a user-centered perspective. Programmers, deep in thoughts of algorithms, “design” user interfaces accidentally or not at all.
Many programmers embrace integrating users frequently (weekly or daily) into the programming process. Although this shares design responsibility with the user, it ignores a serious methodological flaw: a confusion of domain knowledge with design knowledge. Users can articulate problems with an interaction but are not often capable of visualizing solutions. Design is a specialized skill, just like programming. Programmers would never ask users to help them code; design problems should be treated no differently.
Evolution of Software Development Process
Originally programmers did it all In the early days, smart programmers dreamed up useful software, wrote it, and tested it on their own.
Managers brought order Professional managers were brought in. Good product managers understand the market and competitors and define the software product by creating requirements documents. However, requirements are often little more than a list of features, and managers find themselves having to give up features to meet schedule.
Testing and design became separate steps As the industry matured, testing became a separate discipline and a separate step. In the move from command-line to graphical user interface (GUI), design and usability also became involved, though often only at the end, and often only affecting visual presentation. Today common practice includes simultaneous coding and design followed by bug and user testing, then revision.
Design must precede the programming effort A goal-directed design approach means that all decisions proceed from a formal definition of the user and their goals. Definition of the user and user goals is the responsibility of the designer, thus design must precede programming.
Design
Design, according to industrial designer Victor Papanek, is the conscious and intuitive effort to impose meaningful order. Cooper proposes a more detailed definition: • Understanding the user’s wants, needs, motivations, and contexts • Understanding business, technical, and domain requirements and constraints • Translating this knowledge into plans for artifacts whose form, content, and behavior is useful, usable, and desirable, as well as economically viable and technically feasible.
This definition applies across all design disciplines. When performed using appropriate methods, design can provide the missing human connection in technological products.
Three Dimensions of Design & Interaction Design
Interaction design focuses on an area traditional design disciplines do not often explore: the design of behavior. It is only with interactive technologies—courtesy of the computer—that designing the behavior of artifacts and how this behavior affects and supports human goals has become a discipline worthy of attention.
One way to understand the difference is through a historical lens:
- In the first half of the 20th century, designers focused primarily on form.
- Later designers became increasingly concerned with meaning (e.g., retro forms in the 70s).
- Today, information designers include the design of usable content.
Within the last fifteen years, a growing group of designers have begun to talk about behavior: the dynamic ways that software-enabled products interact directly with users. Interactive products must have form, meaning, and behavior in some measure.
🔑 Definition — Interaction Design: Simply put, interaction design is the definition and design of the behavior of artifacts, environments, and systems, as well as the formal elements that communicate that behavior. Unlike traditional design disciplines focused on form and meaning, interaction design seeks first to plan and describe how things behave, and then to describe the most effective form to communicate those behaviors.
⭐ Key Takeaways
The most critical lesson from this lecture is that technical excellence alone guarantees nothing; a product must provide excellent user experience or it will fail in the market. Usability is a core dimension of software quality, and it must be treated strategically (designing how users interact) and tactically (applying specific interface guidelines). The industry’s three main failures are ignorance of users, the conflict of interest where programmers design interfaces, and the lack of a reliable user-centered design process. The solution is to separate design from programming and ensure design precedes coding, using a goal-directed approach. Finally, interaction design is a distinct discipline focused on behavior, not just form or meaning, and is essential for creating products that truly support human goals.
🧠 Quick Revision Questions
- What are the three main reasons why the digital technology industry fails to produce successful, user-friendly products?
- According to the lecture, what is the fundamental flaw in having programmers design user interfaces while also coding them?
- What is the key difference between the strategic and tactical aspects of usability in interface design?
- How does the evolution of the software development process explain the need for design to precede programming?
- Define interaction design and explain how it differs from traditional design disciplines that focus on form and meaning.
📘 Lecture 17 — HCI Process and Methodologies
📖 Overview: This lecture explores the design process in Human Computer Interaction and introduces various lifecycle models that structure software development. Understanding these models is crucial because they determine how user requirements are gathered, how design evolves, and how usability is evaluated throughout a project’s lifecycle.
🗂️ Topics Covered
The lecture begins with multiple definitions of design from different perspectives, then introduces the concept of lifecycle models for software development. It covers the traditional Waterfall model and its flaws, followed by the Spiral lifecycle model with risk analysis and prototyping. Rapid Application Development (RAD) and DSDM are discussed as user-centered approaches. The lecture then presents HCI-specific models including the Star lifecycle model and the Usability Engineering lifecycle, concluding with the Goal-Directed Design process that bridges the gap between research and design.
📝 Lecture Summary
Design Definitions
The term design, even used in the context of designing a computer system, can have different meanings. According to Jones (1981), design can be defined as: ‘Finding the right physical components of a physical structure’, ‘A goal-directed problem-solving activity’, or ‘The imaginative jump from present facts to future possibilities’. On engineering design, Jones found that it involves using scientific principles, technical information, and imagination to define a system that performs pre-specified functions with maximum economy and efficiency. Webster (1988) stresses the relationship between design representation and design process, stating that a design is an information base that describes elaborations of representations.
Design refers to both the process of developing a product and to the various representations produced during that process. Designers need to understand user requirements and represent this understanding in different ways at different stages. Selecting suitable representations is important for exploring, testing, recording, and communicating design ideas within the design team and with users. Two major activities must be undertaken: understanding the requirements of the product and developing the product.
17.1 Lifecycle models
The term lifecycle model represents a model that captures a set of activities and how they are related. Sophisticated models incorporate a description of when and how to move from one activity to the next and a description of the deliverables for each activity. These models allow developers and managers to get an overall view of development so progress can be tracked, deliverables specified, resources allocated, and targets set. Existing models have varying levels of sophistication — for small projects with experienced developers a simple process is adequate, but for larger systems involving tens or hundreds of developers with thousands of users, more formality and discipline is needed.
The waterfall lifecycle was the first model generally known in software engineering and forms the basis of many lifecycles in use today. This is basically a linear model in which each step must be completed before the next step can be started. For example, requirements analysis must be completed before design begins. The lifecycle starts with requirements analysis, moves into design, coding, implementation, testing, and finally maintenance.
🔑 Definition — Waterfall Model: A linear sequential software development lifecycle model where each phase (requirements analysis, design, code, test, maintenance) must be completed before the next phase can begin.
Flaws of waterfall model: One main flaw is that requirements change over time as businesses and the environment change rapidly, so it does not make sense to freeze requirements for months or years while design and implementation are completed. Some feedback to earlier stages was acknowledged as desirable, but the opportunity to review and evaluate with users was not built into this model.
The spiral lifecycle model was suggested by Barry Boehm in 1988. Two features are immediately clear from the model: risk analysis and prototyping. The spiral model incorporates them in an iterative framework that allows ideas and progress to be repeatedly checked and evaluated. Each iteration around the spiral may be based on a different lifecycle model and may have different activities.
🔑 Definition — Spiral Model: An iterative software development lifecycle model that incorporates risk analysis and prototyping, where each iteration may be based on a different lifecycle model and development is driven by plans focused on risks rather than just intended functionality.
💡 Why this matters: Unlike the waterfall, the spiral explicitly encourages alternatives to be considered and problems to be re-addressed. A more recent version called the Win Win spiral model incorporates identification of key stakeholders and their “win” conditions, including a period of stakeholder negotiation to ensure a “win-win” result.
Rapid Application Development (RAD) attempts to take a user-centered view and minimize the risk caused by requirements changing during the project. Two key features of a RAD project are: (1) Time-limited cycles of approximately six months (time-boxing), at the end of which a system or partial system must be delivered, breaking down large projects into smaller projects; and (2) JAD (Joint Application Development) workshops where users and developers come together to thrash out requirements through intensive requirements-gathering sessions with representatives from each stakeholder group.
A basic RAD lifecycle has five phases: project set-up, JAD workshops, iterative design and build, engineer and test final prototype, and implementation review. The popularity of RAD led to DSDM (Dynamic Systems Development Method), developed by a non-profit consortium. The first of nine underlying principles is that “active user involvement is imperative.” The DSDM lifecycle involves five phases: feasibility study, business study, functional model iteration, design and build iteration, and implementation.
17.2 Lifecycle models in HCI
Fewer lifecycle models have arisen from HCI than from software engineering, and they have a stronger tradition of user focus. The first model discussed is the Star model, derived from empirical work on understanding how designers tackled HCI design problems. The second is the usability engineering lifecycle, which shows a more structured approach.
The Star Lifecycle model was proposed by Hartson and Hix in 1989. It emerged from empirical work looking at how interface designers worked, identifying two modes of activity: analytic mode (top-down, organizing, judicial, formal, working from systems view towards user’s view) and synthetic mode (bottom-up, free-thinking, creative, adhoc, working from user’s view towards systems view).
🔑 Definition — Star Lifecycle Model: An HCI lifecycle model where activities are highly interconnected without specifying any ordering, and evaluation is central — you can move from any activity to any other provided you first go through evaluation.
Unlike other lifecycle models, the Star does not specify any ordering of activities. The activities are highly interconnected, and whenever an activity is completed, its results must be evaluated. A project may start with requirements gathering, evaluating an existing situation, analyzing existing tasks, or any other activity.
The Usability engineering lifecycle was proposed by Deborah Mayhew in 1999. Her lifecycle provides a holistic view of usability engineering and a detailed description of how to perform usability tasks, specifying how usability tasks can be integrated into traditional software development lifecycles. It is particularly helpful for those with little expertise in usability.
🔑 Definition — Usability Engineering Lifecycle: A structured HCI lifecycle proposed by Mayhew that has three main tasks (requirements analysis, design/testing/development, and installation) with usability goals captured in a style guide used throughout the project.
The lifecycle has essentially three tasks: requirements analysis, design/testing/development, and installation, with the middle stage being the largest involving many subtasks. It includes identifying requirements, designing, evaluating, and building prototypes. It explicitly includes the style guide as a mechanism for capturing and disseminating usability goals. Some sub-steps can be skipped if unnecessarily complex for the system being developed.
The Goal-Directed Design Process: Most technology-focused companies don’t have an adequate process for user-centered design. Even enlightened organizations face critical issues from traditional approaches. Quantitative market research and market segmentation falls short of providing critical information about how people actually use products. A second problem is that most traditional methods don’t provide a means of translating research results into design solutions — design remains a “black box.”
Bridging the gap: The role of design in the development process needs to change.
Design as product definition: When properly deployed, design identifies user requirements and defines a detailed plan for the behavior and appearance of products. Design provides true product definition based on goals of users, needs of business, and constraints of technology.
Designers as researchers: If design is to become product definition, designers need to take on broader roles. One problem with current development is overspecialization — researchers perform research and designers perform design. What is missing is a systematic means of translating and synthesizing research into design solutions. One powerful tool designers bring is empathy — the ability to feel what others are feeling. Direct exposure to users immerses designers in the users’ world and gets them thinking about users long before proposing solutions. Isolating designers from users eliminates empathic knowledge.
Between research and design: models, requirements, and frameworks: Few design methods incorporate a means of effectively translating research knowledge into detailed design specifications. Few methods capture user behaviors in a manner that appropriately directs product definition — most provide information at the task level rather than about user goals. Instead, explicit systematic processes are needed for defining user models, establishing design requirements, and translating those into a high-level interaction framework.
Goal-Directed Design seeks to bridge the gap between user research and design through a combination of new techniques and known methods brought together in more effective ways. The process flow is: Research (user and domain) → Modeling (users and use context) → Requirements (definition of user, business, and technical needs) → Framework (definition of design structure and flow) → Refinement (of behavior, form, and content).
⭐ Key Takeaways
A student must understand that different lifecycle models offer different approaches to software development, each with distinct advantages and limitations. The Waterfall model is linear and rigid, failing to accommodate changing requirements or user feedback, while the Spiral model introduces iteration through risk analysis and prototyping. RAD and DSDM emphasize user involvement through time-boxing and JAD workshops. The Star lifecycle places evaluation at its center with no prescribed ordering of activities, making it flexible for HCI design. Finally, Goal-Directed Design bridges the critical gap between user research and design solutions by having designers act as researchers and using systematic processes for defining user models, requirements, and frameworks.
🧠 Quick Revision Questions
- What are the five phases of the traditional Waterfall lifecycle model, and what is its main flaw regarding requirements?
- How does the Spiral lifecycle model differ from the Waterfall model, and what two features does it explicitly incorporate?
- What are the two key features of Rapid Application Development (RAD), and what does DSDM stand for?
- In the Star lifecycle model, what is the role of evaluation, and how are activities related to each other?
- What are the five stages of the Goal-Directed Design process, and why is it important for designers to act as researchers?
📘 Lecture 18 — Goal-Directed Design Methodologies
📖 Overview: This lecture introduces the Goal-Directed Design approach, a methodology that balances user desires, business viability, and technical feasibility to create successful interactive products. It emphasizes understanding user goals and behaviors through systematic phases, while also addressing the critical distinction between beginner, intermediate, and expert users.
🗂️ Topics Covered
The lecture covers the Goal-Directed Design model which balances business, engineering, and user concerns across five phases: Research, Modeling, Requirements Definition, Framework Definition, and Refinement. It describes how ethnographic research and user interviews inform the creation of personas and usage patterns, then explains the importance of designing primarily for perpetual intermediate users rather than beginners or experts.
📝 Lecture Summary
18.1 Goal-Directed Design Model
Underlying the goal-directed approach to design is the premise that a product must balance business and engineering concerns with user concerns. You begin by asking, “What do people desire?” then you ask, “Of the things people desire, what will sustain a business.” And finally you ask, “Of the things people desire, that will also sustain the business, what can we build?” A common trap is to focus on technology while losing sight of viability and desirability.
Understanding the importance of each dimension is only the beginning; understanding must also be acted upon. The goal-directed design process is an analog to business and engineering planning processes—it results in a solid user model and a comprehensive interaction plan. The user plan determines the probability that a customer will adopt a product, the business plan determines the probability that the business can sustain itself, and the technology plan determines the probability that the product can be made to work. Multiplying these three factors determines the overall probability that a product will be successful.
💡 Why this matters: This three-dimensional balance prevents designers from building technically impressive products that nobody wants or businesses that cannot sustain themselves.
18.2 A Process Overview
Goal-Directed Design combines techniques of ethnography, stakeholder interviews, market research, product/literature reviews, detailed user models, scenario-based design, and a core set of interaction principles and patterns. It provides solutions that meet the needs and goals of users while also addressing business/organizational and technical imperatives. The process is divided into five phases:
Research Phase: This phase employs ethnographic field study techniques (observation and contextual interviews) to provide qualitative data about potential and/or actual users of the product. It also includes competitive product audits, reviews of market research and technology white papers, and one-on-one interviews with stakeholders, developers, subject matter experts (SMEs), and technology experts. One of the principal outcomes is an emergent set of usage patterns—identifiable behaviors that help categorize modes of use of a potential or existing product. These patterns suggest goals and motivations (specific and general desired outcomes of using the product). In business and technical domains, these behavior patterns tend to map to professional roles; for consumer products, they tend to correspond to lifestyle choices.
Modeling Phase: During the modeling phase, usage and workflow patterns discovered through analysis are synthesized into domain models (information flow and workflow diagrams) and user models or personas—detailed composite user archetypes that represent distinct groupings of behavior patterns, goals, and motivations. Personas serve as the main characters in narrative scenario-based design and represent a powerful communication tool that helps developers and managers understand design rationale and prioritize features based on user needs. Specific design targets are chosen through a process of comparing goals and assigning a hierarchy of priority.
Possible user persona type designations include:
- Primary: the persona's needs are sufficiently unique to require a distinct interface form and behavior
- Secondary: primary interface serves the needs of the persona with a minor modification or addition
- Supplement: the persona's needs are fully satisfied by a primary interface
- Served: the persona is not an actual user of the product, but is indirectly affected by it and its use
- Negative: the persona is created as an explicit, rhetorical example of whom not to design for
Requirements Definition Phase: This phase employs scenario-based design methods focused on meeting the goals and needs of specific user personas. For each interface/primary persona, the process involves an analysis of persona data and functional needs, prioritized and informed by persona goals, behaviors, and interactions with other personas. The analysis is accomplished through iteratively refined context scenarios that start with a “day in the life” of the persona using the product, describing high-level product touch points, and thereafter successively defining detail at ever-deepening levels. The output is a requirements definition that balances user, business, and technical requirements.
Framework Definition Phase: Teams synthesize an interaction framework by employing general interaction design principles and a set of interaction design patterns that encode general solutions to classes of previously analyzed problems. These patterns are hierarchically organized and continuously evolve. After data and functional needs are described at a high level, they are translated into design elements according to interaction principles, then organized into design sketches and behavior descriptions. The output is an interaction framework definition—a stable design concept that provides the logical and gross formal structure for the detail to come.
Refinement Phase: The refinement phase proceeds similarly to the framework definition phase but with greater focus on task coherence, using key path and validation scenarios focused on storyboarding paths through the interface in high detail. The culmination is detailed documentation of the design—a form and behavior specification delivered in either paper or interactive media.
Goals, Not Features, Are Key to Product Success
Programmers and engineers share a strong tendency to think about products in terms of functions and features because this is how developers build software: function-by-function. The problem is that this is not how users want to use it. The deciding factor for whether a feature should be included should never be simply that we have the technical capability to do it; instead, the deciding factor should be whether that feature directly or indirectly helps to achieve the goals of the user while still meeting the needs of the business.
The successful interaction designer must be sensitive to user's goals amid the pressures and chaos of the product-development cycle. The Goal-Directed process, with its clear rationale for design decisions, makes persuading engineers easier, keeps marketing and management stakeholders in the loop, and ensures that the design isn't just guesswork or a reflection of team members' personal preferences.
18.3 Types of Users
Most users are neither beginners nor experts; instead they are intermediates. If we graph number of people against skill level, a relatively small number of beginners are on the left side, a few experts are on the right, and the majority—intermediate users—are in the center. The bell curve is a snapshot in time. Although beginners do not remain beginners for very long, and experts come and go rapidly, beginners change even more rapidly. Both beginners and experts tend over time to gravitate towards intermediacy.
Nobody remains a beginner for long because people don't like to be incompetent. Beginners become intermediates very quickly—or they drop out altogether. Most users thus remain in a perpetual state of adequacy striving for fluency, with their skills ebbing and flowing depending on how frequently they use the program. Larry Constantine first identified the importance of designing for improving intermediates, but the term perpetual intermediates is preferred because although beginners quickly improve to become intermediates, they seldom go on to become experts.
A well-balanced user interface devotes the bulk of its efforts to satisfying the perpetual intermediate. At the same time, it avoids offending either beginners or experts, recognizing that both are vital. Most users in this middle state would like to learn more about the program but usually don't have the time. Sometimes they use the product extensively and learn new things; other times they don't use the program for months and forget significant portions of what they knew.
Optimizing for Intermediates: Programmers qualify as experts in the software they code because they must explore every possible use case, so their natural tendency is to design implementation model software with every possible option given equal emphasis. Meanwhile, sales, marketing, and management—who demonstrate products to beginners—lobby for bending the interface to serve beginners. Programmers create interaction suitable only for experts, marketers demand interactions suitable only for beginners, but the largest, most stable, and most important group of users is the intermediate group.
Our goal is threefold: to rapidly and painlessly get beginners into intermediacy; to avoid putting obstacles in the way of those intermediates who want to become experts; and most of all, to keep perpetual intermediates happy as they stay firmly in the middle of the skill spectrum.
What Perpetual Intermediates Need: They need access to tools, not explanations of scope and purpose. Tooltips are the perfect perpetual intermediate idiom—they state function in the briefest of terms, consuming minimal video space. Perpetual intermediates know how to use reference materials, so online help via a comprehensive index is a perpetual intermediate tool. They will establish a working set of frequently used functions, and these tools must be placed front-and-center in the user interface. Perpetual intermediates also find it reassuring to know that advanced features exist, even if they don't need them, as it confirms they made the right choice in investing in the program.
⭐ Key Takeaways
The Goal-Directed Design process consists of five sequential phases—Research, Modeling, Requirements Definition, Framework Definition, and Refinement—that balance user, business, and technical concerns. Personas are the central design tool, with primary personas dictating distinct interface forms and other persona types (secondary, supplement, served, negative) having lesser influence. The most critical insight is that most users are perpetual intermediates, not beginners or experts, so design should prioritize making beginners rapidly become intermediates while keeping the majority happy. Programmers tend to design for experts, while marketers push for beginners, yet both approaches ignore the largest user segment. The key to product success is focusing on user goals rather than features.
🧠 Quick Revision Questions
- What are the five phases of the Goal-Directed Design process, and what is the primary output of each phase?
- How does the Goal-Directed Design model balance user desires, business viability, and technical feasibility?
- What are the five persona type designations, and how does each influence the design?
- Why are most users considered "perpetual intermediates," and what should designers do to accommodate them?
- What is the difference between designing for features versus designing for goals, and why is this distinction important?
📘 Lecture 19 — User Research Part-I
📖 Overview: This lecture introduces the foundational concepts of user research in Human Computer Interaction, focusing on the critical distinction between qualitative and quantitative research methods. It emphasizes why understanding users through qualitative techniques is essential for successful design, and provides a detailed exploration of various qualitative research methods that help designers gather rich, contextual information about users, their behaviors, and their environments.
🗂️ Topics Covered
The lecture begins by contrasting qualitative and quantitative research, explaining why human behavior is too complex for purely numerical analysis. It then details the value of qualitative research for understanding user domains, contexts, and constraints. The main body covers six types of qualitative research techniques: stakeholder interviews, subject matter expert interviews, user and customer interviews, user observation/ethnographic field studies, literature review, and product/prototype and competitive audits, each with specific guidance on when and how to apply them.
📝 Lecture Summary
Qualitative versus Quantitative Research
Research is often associated with science and objectivity, but this biases people toward thinking only quantitative data is valid. The notion that "numbers don’t lie" is prevalent, though numbers ascribed to human activities can be manipulated or reinterpreted as dramatically as words. Data from hard sciences like physics is fundamentally different from data on human activities—electrons don’t have moods that vary minute to minute, and the tight controls physicists use are impossible in social sciences. Reducing human behavior to statistics overlooks important nuances that make an enormous difference to product design. Quantitative research can only answer questions about how much or how many along a few reductive axes, while qualitative research tells you about what, how and why in rich, multivariate detail. Social scientists have long recognized that human behaviors are too complex and subject to too many variables to rely solely on quantitative data.
🔑 Definition — Quantitative Research: A type of research that yields numerical data, answering questions about "how much" or "how many" along reductive axes. 🔑 Definition — Qualitative Research: A type of research that provides rich, multivariate detail about what, how, and why regarding human behaviors and experiences. 💡 Why this matters: Understanding this distinction is crucial because designing for human users requires understanding context and meaning, not just countable metrics.
The value of qualitative research
Qualitative research helps understand the domain, context and constraints of a product in more useful ways than quantitative research. It quickly helps identify patterns of behavior among users and potential users. Specifically, qualitative research helps understand: existing products and how they are used; potential users and how they currently approach activities and problems; technical, business, and environmental contexts of the product; and vocabulary and other social aspects of the domain. It also helps design projects by providing credibility and authority to the design team, uniting the team with a common understanding, and empowering management to make more informed decisions. Qualitative methods tend to be faster, less expensive, more flexible, and more likely to provide useful answers to important design questions such as: what problems are people encountering? Into what broader contexts does the product fit? What are the basic goals and tasks?
📌 Example: Qualitative research can answer "What problems are people encountering with their current ways of doing what the product hopes to do?"—a question that quantitative metrics cannot adequately address.
19.1 Types of qualitative research
Social science and usability texts contain many methods for conducting qualitative research. The following techniques are discussed: stakeholder interviews, subject matter expert (SME) interviews, user and customer interviews, user observation/ethnographic field studies, literature review, and product/prototype and competitive audits.
Stakeholder interviews
Research for any new product should start by understanding the business and technical context in which the product will be built. Stakeholders are any key members of the organization commissioning the design work, typically including managers and key contributors from engineering, sales, product marketing, marketing communications, customer support, and usability. They may also include similar people from partner organizations and executives. Interviews with stakeholders should occur before any user research begins. It is most effective to interview each stakeholder one-on-one to promote candor and ensure individual views are not lost. Interviews need not last longer than about an hour. Important information to gather includes: the preliminary vision of the product from each stakeholder perspective; budget and schedule; technical constraints; business drivers; and stakeholders’ perceptions of the user. Understanding these issues helps designers better serve their customers and build internal consensus critical for decision making.
🔑 Definition — Stakeholders: Any key members of the organization commissioning the design work, including managers and key contributors from various departments. 💡 Why this matters: Different business departments may have slightly different and incomplete perspectives on the product, like the fable of the blind men and the elephant, and these must be harmonized with user perspectives.
Subject matter expert (SME) interviews
Some stakeholders may also be subject matter experts (SMEs) : experts on the domain within which the product will operate. Most SMEs were users of the product or its predecessors at one time and may now be trainers, managers, or consultants. They can provide valuable perspective, but designers should recognize that SMEs represent a somewhat skewed perspective. Points to consider: SMEs are expert users whose long experience may have made them accustomed to current interactions; they are not designers, so the most useful information from their suggestions is the causative problems leading to their proposed solutions; SMEs are necessary in complex or specialized domains such as medical, scientific, or financial services; and you will want access to SMEs throughout the design process for reality checks on design details.
🔑 Definition — Subject Matter Expert (SME): An expert on the domain within which the product will operate, often a former user now serving as trainer, manager, or consultant.
User and customer interviews
It is easy to confuse users with customers. For consumer products, customers are often the same as users, but in corporate or technical domains, they rarely describe the same people. Customers are those who make the decision to purchase the product. For consumer products, customers are frequently users, though for products aimed at children, customers are parents. In enterprise products, the customer is often an IT manager with distinct goals. It's important to understand customers and their goals to make a product viable, while realizing that customers seldom use the product themselves. When interviewing customers, you want to understand: their goals in purchasing, frustrations with current solutions, decision process, role in installation and management, and domain issues. Users should be the main focus of the design effort—they are the people personally trying to accomplish something with the product. A good set of user interviews includes both current and potential users. Information from users includes: problems and frustrations, context of use, domain knowledge from a user perspective, current tasks, and user goals.
🔑 Definition — Customers: People who make the decision to purchase a product, who may or may not be the actual users. 🔑 Definition — Users: People who personally try to accomplish something with the product, should be the main focus of the design effort.
User observation
Most people cannot accurately assess their own behaviors, especially outside the context of their activities. Interviews performed outside the context of situations yield less complete and less accurate data. You can talk to users about how they think they behave, or observe it first hand—the latter provides superior results. Many usability professionals use technological aids like audio or video recorders, but care must be taken not to make these too obtrusive. Perhaps the most effective technique for gathering qualitative user data combines interviews and observation, allowing the designer to ask clarifying questions about situations and behaviors observed in real-time.
Literature review
In parallel with stakeholder interviews, the design team should review any literature pertaining to the product or its domain. This can include product marketing plans, market research, technology specifications and white papers, business and technical journal articles, competitive studies, Web searches, usability study results, and customer support data. The design team should collect this literature, use it as a basis for developing questions, and later use it to supply additional domain knowledge and vocabulary.
Product and competitive audits
Also in parallel to stakeholder and SME interviews, it is helpful for the design team to examine any existing version or prototype of the product, as well as its chief competitors. Doing so gives the design team a sense of the state of the art and provides fuel for questions during interviews. The team should engage in an informal heuristic or expert review of both the current and competitive interfaces, comparing each against interaction and visual design principles. This familiarize the team with strengths and limitations of what is currently available.
🔑 Definition — Heuristic Review: An informal evaluation of a user interface against established usability principles or "heuristics."
⭐ Key Takeaways
The most critical distinction from this lecture is between qualitative research, which explores what, how, and why in rich detail, and quantitative research, which only measures how much or how many—human behavior is too complex for purely numerical analysis. Students must remember that qualitative research is essential for understanding domain context, user behaviors, and design constraints, and it tends to be faster and more flexible than quantitative approaches. The six types of qualitative research each serve a specific purpose: stakeholder interviews provide business and technical context; SME interviews offer specialized domain expertise; user and customer interviews reveal distinct perspectives from purchasers versus actual users; user observation captures authentic behavior that self-reporting cannot; literature review supplies background knowledge; and competitive audits establish the state of the art. A critical point is that customers (who buy the product) and users (who use it) are often different people, especially in enterprise contexts, and both must be understood separately. Finally, the most effective user research combines interviews with direct observation in the user's natural context, as people cannot accurately report their own behaviors outside their environment.
🧠 Quick Revision Questions
- What are the fundamental differences between qualitative and quantitative research in terms of the types of questions they can answer?
- Why is it important to conduct stakeholder interviews before any user research begins?
- What is the key distinction between customers and users, and why must designers treat them differently?
- Why does the lecture claim that user observation provides superior results compared to interviews alone?
- What types of information should a design team gather during a literature review, and how should this information be used?
📘 Lecture 20 — User Research Part-II
📖 Overview: This lecture continues the study of qualitative research techniques, focusing specifically on the ethnographic field study method. It explores the user-centered design philosophy and provides detailed guidance on conducting ethnographic interviews, including preparation, frameworks, and practical implementation strategies for gathering rich user data in natural contexts.
🗂️ Topics Covered
The lecture covers the user-centered approach and its three core principles, ethnographic field study methods including the three dimensions framework (distributed coordination, plans and procedures, awareness of work), contextual design and contextual inquiry techniques, and the comprehensive process of preparing for ethnographic interviews including persona hypothesis development, identifying candidates, and creating interview plans.
📝 Lecture Summary
20.1 User-Centered Approach
The user-centered approach means that real users and their goals, not just technology, should be the driving force behind product development. A well-designed system should make the most of human skill and judgment, be directly relevant to the work at hand, and support rather than constrain the user. This is less a technique and more a philosophy.
In 1985, Gould and Lewis laid down three principles they believed would lead to a "useful and easy to use computer system":
- Early focus on users and tasks: Understand who the users will be by directly studying their cognitive, behavioral, anthropomorphic, and attitudinal characteristics through observing users doing normal tasks and studying the nature of those tasks.
- Empirical measurement: Early in development, observe and measure the reactions and performance of intended users to printed scenarios and manuals. Later, users interact with simulations and prototypes while their performance and reactions are observed, recorded, and analyzed.
- Iterative design: When problems are found in user testing, fix them and carry out more tests and observations. Design and development cycles of "design, test, measure, and redesign" are repeated as often as necessary.
Applying ethnography in design
Ethnography is a method that comes originally from anthropology and literally means "writing the culture." It displays the social organization of activities to understand work. It aims to find the order within an activity rather than impose any framework of interpretation. Observers immerse themselves in the users' environment, participate in day-to-day work, join conversations, attend meetings, and read documents. The aim is to make the implicit explicit.
🔑 Definition — Ethnography: A broad-based approach where users are observed as they go about their normal activities, with observers immersing themselves in the users' environment to understand the social organization of work and make implicit practices explicit.
Beynon-Davies suggested ethnography can be associated with development as:
- Ethnography of development: Studies of developers themselves and their workplace
- Ethnography for development: Studies that can be used as a resource for development
- Ethnography within software development: Techniques integrated into methods for development
💡 Why this matters: The ethnographic experience reveals information missed by other methods, such as how people do "real" work versus formal procedures, the nature of collaboration, awareness of others' work, and implicit goals not recognized by workers themselves.
📌 Example: Heath et al. studied dealers in a stock exchange to see whether proposed technological support for market trading was suitable. They observed the process of writing tickets to record deals. While others suggested introducing touch screens and headphones, Heath et al. discovered these proposals were misguided — touch screens would reduce information availability to others, and headphones would impede dealers' ability to monitor one another. They recommended pen-based mobile systems with gesture recognition instead.
"In some way, the goals of design and the goals of ethnography are at opposite ends of a spectrum. Design is concerned with abstraction and rationalization. Ethnography, on the other hand, is about detail."
20.2 Ethnography Framework
The ethnographic framework has been developed specifically to help structure the presentation of ethnographies in a way that enables designers to use them. This framework has three dimensions:
1. Distributed coordination Focuses on the distributed nature of tasks and activities, and the means and mechanisms by which they are coordinated. This has implications for the kind of automated support required.
2. Plans and procedures Focuses on organizational support for work, such as workflow models and organizational charts, and how these support the work. Understanding this impacts how the system is designed to utilize this kind of support.
3. Awareness of work Focuses on how people keep themselves aware of others' work. Being aware of others' actions and work activities can be a crucial element of doing a good job. Implications relate to the sharing of information.
An alternative approach is to train developers to collect ethnographic data themselves, giving designers first-hand experience of the situation. Two methods provide support for this:
Coherence
The coherence method combines experiences of using ethnography to inform design with developments in requirements engineering. It is intended to integrate social analysis with object-oriented analysis from software engineering.
🔑 Definition — Viewpoints: Within Coherence, these are focus questions for each of the three framework dimensions, intended to guide the observer to particular aspects of the workplace.
🔑 Definition — Concerns: A kind of goal representing criteria that guide the requirements activity. These are addressed within each appropriate viewpoint. The concerns are:
- Paper work and computer work: Embodiments of plans and procedures, and mechanisms for developing and sharing awareness of work
- Skill and the use of local knowledge: The "workarounds" developed in organizations that are at the heart of how real work gets done
- Spatial and temporal organization: The physical layout of the workplace and areas where time is important
- Organizational memory: Formal documents are not the only way things are remembered — individuals may keep their own records, or there may be local gurus
Contextual design
Contextual design was developed to handle the collection and interpretation of data from fieldwork with the intention of building a software-based product. It provides a structured approach to gathering and representing information from fieldwork.
Contextual design has seven parts:
- Contextual inquiry
- Work modeling, consolidation
- Work redesign
- User environment design
- Mockup
- Test with customers
- Putting it into practice
Contextual inquiry
Contextual inquiry is based on a master-apprentice model of learning: observing and asking questions of the users as if she is the master craftsman and he interviews the new apprentice. Four basic principles for engaging in ethnographic interview:
- Context: Interact with and observe the user in their normal work environment, filled with the artifacts they use each day
- Partnership: The interview should take the tone of a collaborative exploration, alternating between observation and discussion of its structure and details
- Interpretation: Much of the designer's work is reading between the lines of facts gathered about user's behaviors, their environment, and what they say — but avoid assumptions without verifying with users
- Focus: The designer needs to subtly direct the interview to capture data relevant to design issues
Improving on contextual inquiry
Process improvements for a more highly leveraged research phase:
- Shortening the interview process: Interviews as short as one hour are sufficient with about six well-selected users for each hypothesized role
- Using smaller design teams: More effective to conduct interviews sequentially with the same designers (two or three)
- Identifying goals first: Ethnographic interviews should first identify and prioritize user goals before determining tasks
- Looking beyond business contexts: Ethnographic interviews are also possible in consumer domains
20.3 Preparing for Ethnographic Interviews
Ethnographic interviews take the spirit of anthropological research and apply it on a micro level — understanding the behaviors and rituals of people interacting with individual products.
Identifying candidates
Designers must capture an entire range of user behaviors regarding a product. Based on information from stakeholders, SMEs, and literature reviews, designers create a hypothesis that serves as a starting point for determining what sorts of users to interview.
Kim Goodwin coined this the persona hypothesis, the first step towards identifying and synthesizing personas. The persona hypothesis is based on likely behavioral differences, not demographics, but takes into consideration identified target markets and demographics.
🔑 Definition — Persona hypothesis: A first cut at defining the different kinds of users (and sometimes customers) for a product in a particular domain. It serves as a basis for an initial set of interviews and attempts to address three questions:
- What different sorts of people might use this product?
- How might their needs and behaviors vary?
- What ranges of behavior and types of environments need to be explored?
Roles in business and customer domains
For business products, roles — common sets of tasks and information needs related to distinct classes of users — provide an important initial organizing principle. For example, in an enterprise portal: people who search for content, people who upload and update content, and people who technically administer the portal.
Unlike business users, consumers don't have concrete job descriptions. Their roles map more closely to lifestyle choices, and it is possible for consumer users to assume multiple roles even for a single product.
Behavioral and demographic variables
Beyond roles, a persona hypothesis seeks to identify behavioral variables that might distinguish users based on their needs and behaviors. Examples: frequency of shopping, desire to shop, motivation to shop.
Another helpful approach is making use of demographic variables such as ages, locations, gender, and incomes of target markets.
Domain expertise versus technical expertise
Important behavioral distinction: technical expertise (knowledge of digital technology) versus domain expertise (knowledge of a specialized subject area pertaining to a product). Domain support may be a necessary part of the product's design, as well as technical ease of use.
Environmental variables
Cultural differences between organizations, especially for business products. Examples: company size (small to multinational), IT presence (ad hoc to draconian), security level (lax to tight).
20.4 Putting a Plan Together
After creating a persona hypothesis with potential roles, behavioral, demographic, and environmental variables, create an interview plan that can be communicated to the person in charge of providing access to users.
Each identified role, behavioral variable, demographic variable, and environmental variable should be explored in four to six interviews (sometimes more if a domain is particularly complex). However, interviews can overlap — interviewing a female in her twenties who loves to shop counts for multiple variables. By cleverly mapping variables to interviewee screening profiles, the number of interviews can be kept manageable.
⭐ Key Takeaways
The user-centered approach demands that real users and their goals drive product development through early focus on users, empirical measurement, and iterative design cycles. Ethnography is the most immersive qualitative technique, making implicit workplace behaviors explicit by observing users in their natural context. The three-dimensional ethnographic framework (distributed coordination, plans and procedures, awareness of work) helps structure ethnographic data for design use, while contextual design and coherence provide structured approaches for translating fieldwork into design specifications. Successful ethnographic interviews require careful preparation through persona hypothesis development, identifying diverse user types across behavioral, demographic, and environmental variables, and planning 4-6 interviews per variable with overlapping coverage.
🧠 Quick Revision Questions
- What are the three principles laid down by Gould and Lewis for creating useful and easy-to-use computer systems?
- What is the fundamental difference between the goals of design and the goals of ethnography?
- What were Heath et al.'s findings regarding proposed technological support for stock exchange dealers, and what alternative did they recommend?
- What are the four concerns in the Coherence method, and what does each represent?
- What is a persona hypothesis, and what three questions does it attempt to address?
📘 Lecture 21 — User Research Part-III
📖 Overview: This lecture completes the user research series by focusing on ethnographic interview techniques and other qualitative research methods. It explains how to conduct effective interviews that reveal user goals, behaviors, and workflows, while also introducing alternative research approaches like focus groups and market demographics for comprehensive user understanding.
🗂️ Topics Covered
The lecture covers methodologies for conducting ethnographic interviews including securing interviews, interview team structure, and timing. It details the three phases of ethnographic interviews (early, mid, late-phase) with their respective questioning approaches. Basic interview methods are presented along with four categories of interview questions: goal-oriented, system-oriented, workflow-oriented, and attitude-oriented. The lecture also surveys types of qualitative research, other research types including focus groups and market demographics, and concludes with a comparison table of different research techniques.
📝 Lecture Summary
21.1 How to conducting ethnographic interviews
Securing interviews can be achieved through three primary channels: stakeholders, market or usability research firms, and friends and relatives. These sources help researchers gain access to appropriate interview participants.
For interview teams and timings, the recommended structure involves 2 interviewees per day, with each interview lasting 1 hour, allowing for a maximum of 6 interviews per day when working in pairs.
21.2 Phases of ethnographic interviews
The interview process moves from beginning to end along two dimensions: from structural issues to specific issues, and from goal-oriented issues to task-oriented issues.
Early-phase interviewing is exploratory, focused on gaining domain knowledge through open-ended questions that allow participants to share broad perspectives without constraint.
Mid-phase interviewing shifts toward identifying patterns of use by asking clarifying questions and more focused questions that build on initial findings.
Late-phase interviewing aims to confirm patterns of use and clarify user roles and behaviors through closed-ended questions that validate or challenge emerging hypotheses.
21.3 Basic interview methods
Effective interviews follow several key principles: interview where the action happens (in the user's natural environment), avoid a fixed set of questions (remain flexible), focus on goals first, tasks second (understand intent before actions), avoid making the user a designer (don't ask users to solve design problems), avoid discussions of technology (focus on user experience), encouraging storytelling (narratives reveal context), ask for a show-and-tell (demonstrations provide rich data), and avoid leading questions (don't bias responses).
Goal-oriented questions explore user objectives:
- Opportunity: "What activities currently waste your time?"
- Goals: "What makes a good day? A bad day?"
- Priorities: "What is the most important to you?"
- Information: "What helps you make decisions?"
System-oriented questions examine product interaction:
- Function: "What are the most common things you do with the product?"
- Frequency: "What parts of the product do you use most?"
- Preference: "What are your favorite aspects of the product? What drives you crazy?"
- Failure: "How do you work around problems?"
- Expertise: "What shortcuts do you employ?"
Workflow-oriented questions reveal process and sequence:
- Process: "What did you do when you first came into today? And after that?"
- Occurrence and recurrence: "How often do you do this? What things do you do weekly, monthly but not every day?"
- Exceptions: "What constitutes a typical day? What would be an unusual event?"
Attitude-oriented questions probe emotions and motivations:
- Aspiration: "What do you see yourself doing five years from now?"
- Avoidance: "What would you prefer not to do? What do you procrastinate on?"
- Motivation: "What do you enjoy most about your job (or lifestyle)? What do you always tackle first?"
💡 Why this matters: These question categories systematically uncover different layers of user experience—from surface-level tasks to deep-seated goals and attitudes—ensuring comprehensive understanding.
21.4 Types of qualitative research
The lecture identifies several types of qualitative research:
- Stakeholder interview — interviewing people with vested interest in the product
- Subject matter experts (SME) interviews — gathering specialized domain knowledge
- User and customer interviews — direct feedback from end users
- Literature review — studying existing documentation and research
- Product/prototype and competitive audits — examining existing solutions
- User observation/ethnographic field studies — watching users in natural settings
21.5 Others types of research
Focus groups are used by marketing organizations for traditional product marketing. In this method, representative users are gathered in a room, shown a product and reactions gauged, with reactions recorded by audio/video. The lecture notes limitations of this approach (such as group dynamics potentially skewing individual opinions).
Market demographics and segments research asks "What motivates people to buy?" This is determined by market segmentation that group people by distinct needs and determines who will be receptive to what marketing message or a particular product. Data includes:
- Demographic data: Race, education, income, location
- Psychographic data: Attitude, lifestyle, values, ideology
Usability and user testing is mentioned as a third type of research, though not elaborated in this lecture.
21.6 Comparison of different techniques
The lecture provides a comparison table evaluating different research techniques:
Interviews:
- Good for: Exploring issues
- Kind of data: Some quantitative but mostly qualitative data
- Advantages: Interviewer can guide interviewee if necessary; encourages contact between developers and users
- Disadvantages: Time consuming; artificial environment may intimidate interviewee
Studying documentation:
- Good for: Learning about procedures, regulations and standards
- Kind of data: Quantitative
- Advantages: No time commitment from users required
- Disadvantages: Day-to-day working will differ from documented procedures
Naturalistic observation:
- Good for: Understanding context of user activity
- Kind of data: Quantitative
- Advantages: Observing actual work gives insights that other techniques can't give
- Disadvantages: Very time consuming; huge amounts of data
Focus groups and workshops:
- Good for: Collecting multiple viewpoints
- Kind of data: Some quantitative but mostly qualitative data
- Advantages: Highlights areas of consensus and conflict; encourages contact between developers and users
- Disadvantages: Possibility of dominant characters
Modeling Research: Use ethnographic research techniques to obtain qualitative data:
- User observation
- Contextual interviews
Qualitative data derived from this research informs:
- Usage patterns: Sets of observed behaviors that categorize modes of use
- Goals: Specific and general desired outcomes of using the product
- Personas: (implied as the ultimate output of synthesizing this data)
⭐ Key Takeaways
The ethnographic interview process is a structured yet flexible methodology that progresses through three phases—early (exploratory, open-ended), mid (pattern identification, focused), and late (confirmation, closed-ended)—each requiring different questioning approaches. Effective interviews prioritize goals over tasks, avoid making users designers, and use storytelling and show-and-tell techniques to capture rich contextual data. The four categories of interview questions (goal-oriented, system-oriented, workflow-oriented, and attitude-oriented) systematically uncover different dimensions of user experience. Alternative research methods like focus groups and market demographics serve different purposes and have distinct advantages and disadvantages compared to interviews and observation. Ultimately, all ethnographic research aims to produce qualitative data that reveals usage patterns and user goals, which form the foundation for creating personas and designing user-centered products.
🧠 Quick Revision Questions
- What are the three phases of ethnographic interviewing, and what type of questions is each phase characterized by?
- List at least five of the basic interview methods discussed in the lecture (principles for conducting effective interviews).
- What are the four categories of interview questions, and give one example question for each category?
- What are the advantages and disadvantages of using naturalistic observation compared to interviews?
- How does market demographics and segmentation research differ from ethnographic interviews in terms of what it measures and why?
📘 Lecture 22 — User Modeling
📖 Overview: This lecture introduces the concept of user modeling as a powerful interaction design tool, focusing on personas as composite archetypes based on behavioral data from real users. It explains why models are necessary, how personas overcome common design problems, and provides a detailed framework for constructing and using personas to create user-centered products.
🗂️ Topics Covered
The lecture covers the importance of modeling in design, the concept and strengths of personas as user models, how personas resolve user-centered design issues like the elastic user and self-referential design, the relationship between personas and other user representations, the role of goals in driving behavior including life, experience, and end goals, and a step-by-step process for constructing personas with various types.
📝 Lecture Summary
22.1 Why Model?
Models are powerful tools for representing complex structures and relationships, helping designers make sense of unstructured raw data. Good models emphasize salient features and de-emphasize less significant details. Because we design for users, we must understand and visualize their relationships with each other, their social and physical environment, and the products we design.
Just as physicists create models of the atom from observed data, designers create models of users based on observed behaviors and intuitive synthesis of patterns in the data. Personas provide this formalization, allowing designers to systematically construct patterns of interaction that match user behaviors, mental models, and goals.
💡 Why this matters: Without models, designers have no organizing principle for user data, leading to unfocused design decisions.
22.2 Personas
To create a product for a broad audience, logic suggests making it as broad in functionality as possible. This logic is flawed. The best way to accommodate a variety of users is to design for specific types of individuals with specific needs. Broadly extending functionality increases cognitive load and navigational overhead for all users.
📌 Example: Designing an automobile that pleases every possible driver results in a car with every feature that pleases nobody. Designing different cars for different people with specific goals creates satisfying designs for others with similar needs.
The key is choosing the right individuals whose needs represent a larger set of key constituents and prioritizing design elements to address the most important users without inconveniencing secondary users.
Strengths of personas as a design tool:
- Determine what a product should do and how it should behave
- Communicate with stakeholders, developers, and other designers
- Build consensus and commitment to the design
- Measure the design’s effectiveness
- Contribute to other product-related efforts like marketing and sales
Personas and user-centered design: Personas resolve three issues:
The elastic user: The term "user" is imprecise—every team member has their own conception. This "user" becomes elastic, bending to fit whoever has the floor. Designing for the elastic user gives developers license to code as they please. Real users and personas are not elastic; they have specific requirements based on goals, capabilities, and contexts.
Self-referential design: This occurs when designers project their own goals, motivations, skills, and mental models onto a product. Most "cool" product designs fall into this category. Programmers apply self-referential design when creating implementation-model products that only they understand.
Design edge cases: Personas help prevent designing for edge cases—situations that might happen but usually won't for target personas. Edge cases must be programmed for but should never be the design focus.
Personas are based on research: The primary source of data must be from ethnographic interviews, contextual inquiry, or similar dialogues with actual and potential users. Supporting data includes interviews outside use contexts, stakeholder information, market research, market segmentation models, and literature reviews. However, nothing replaces direct interaction with and observation of users in their native environments.
Personas are represented as individuals: Personas are user models represented as specific, individual humans, synthesized from observations of real people. They engage the empathy of the development team toward the human target of design. Empathy is critical for designers making decisions based on both cognitive and emotional dimensions.
Personas represent classes of users in context: Personas encapsulate a distinct set of usage patterns identified through analysis of ethnographic interviews. They are composite user archetypes assembled by clustering related usage patterns observed across individuals in similar roles.
Personas and reuse: Personas must be context-specific—focused on behaviors and goals related to a specific product domain. They cannot easily be reused across products, even closely linked ones, because behavioral focus may differ.
Archetypes versus stereotypes: Don't confuse persona archetypes with stereotypes. Stereotypes represent designer biases and assumptions, not factual data. Personas developed with inadequate research risk becoming stereotypical caricatures. Personas must be developed and treated with dignity and respect.
Personas explore ranges of behavior: Personas do not seek an average user but identify exemplary types of behaviors along identified ranges. Designers must identify a collection or cast of personas associated with any given product.
Personas must have motivations: All humans have motivations that drive behaviors. Personas capture these motivations in the form of goals. Understanding why a user performs certain tasks gives designers power to improve or eliminate those tasks while still accomplishing the same goals.
Personas versus user roles: User roles are abstractions defining relationships between a class of users and their problems. Problems with user roles include difficulty identifying relationships in the abstract, focus on tasks while neglecting goals, and difficulty bringing models together as a coherent tool. Personas incorporate the same relationships but express them in terms of goals and examples in narrative.
Personas versus user profile: User profiles are often a name attached to brief demographic data with a short fictional paragraph. This is likely a user stereotype and not useful as a design tool. Personas use names and details sparingly as narrative tools to communicate real data.
Personas versus market segments: Market segments are based on demographics and distribution channels; design personas are based on behaviors and goals. Marketing personas shed light on the sales process; design personas shed light on the development process.
User personas versus non-user personas: A frequent error is targeting people who review, purchase, or administer the product but are not end users. Designing for the purchaser is a frequent mistake. For enterprise systems requiring administrators, it's appropriate to create non-user personas with expanded research.
22.3 Goals
If personas provide context for observed behaviors, goals are the drivers behind those behaviors. A persona without goals can serve as a communication tool but is useless as a design tool. User goals serve as a lens through which designers must consider product functions.
Goals motivate usage patterns: Goals provide the answer to why and how personas desire to use a product and serve as shorthand for complex behaviors.
Goals must be inferred from qualitative data: You cannot ask a person directly what their goals are—they either cannot articulate them or won't be accurate. Designers must reconstruct goals from observed behaviors, answers to other questions, non-verbal cues, and environmental clues.
22.4 Types of Goals
Goals come in many varieties. The most important from a user-centered design standpoint are user goals, which are first priority, especially for consumer products.
User goals:
- Life goals
- Experience goals
- End goals
Life goals: Represent personal aspirations that typically go beyond the product context. They represent deep drives and motivations explaining why the user seeks to accomplish end goals.
📌 Examples: Be the best at what I do; Get onto the fast track; Learn all there is to know about this field; Be a paragon of ethics, modesty, and trust.
Life goals rarely figure directly into interface design but are worth keeping in mind.
Experience goals: Simple, universal, personal goals expressing how someone wants to feel while using a product or the quality of their interaction.
📌 Examples: Don't make mistakes; Feel competent and confident; Have fun.
Experience goals represent unconscious goals that people bring to any software product without consciously realizing it.
End goals: Represent the user's expectations of tangible outcomes from using a specific product.
📌 Examples: Find the best price; Finalize the press release; Process the customer's order; Create a numerical model of the business.
End goals must be met for users to think a product is worth their time and money.
Non-user goals:
- Customer goals
- Corporate goals
- Technical goals
Customer goals: Consumer customers often have concerns about safety and happiness. Enterprise customers (IT managers) often have concerns about security, ease of maintenance, and customization.
Corporate goals: Business requirements at a high level.
📌 Examples: Increase profit; Increase market share; Defeat the competition; Use resources more efficiently; Offer more products or services.
Technical goals: Goals that ease the task of software creation, often taking precedence at the expense of users' goals.
📌 Examples: Save money; Run in a browser; Safeguard data integrity; Increase program execution efficiency.
22.5 Constructing Personas
Creating believable and useful personas requires detailed analysis and creative synthesis. The process involves seven steps:
1. Revisit the persona hypothesis: Compare patterns in the data to assumptions made in the persona hypothesis. Were possible roles distinct? Were behavioral variables valid? If data varies from assumptions, add, subtract, or modify roles and behaviors, potentially conducting additional interviews.
2. Map interview subjects to behavioral variables: Map each interviewee against each variable range that applies. The precision is less critical than identifying placement of interviewees in relationship to each other. The way multiple subjects cluster on each variable axis is significant.
3. Identify significant behavior patterns: Look for clusters of subjects occurring across multiple ranges or variables. A set clustering in six to eight variables likely represents a significant behavior pattern forming the basis of a persona. For a pattern to be valid, there must be a logical or causative connection between clustered behaviors.
4. Synthesize characteristics and relevant goals: For each significant pattern, synthesize details describing the potential use environment, typical workday, current solutions, frustrations, and relevant relationships. Use bullet points based on observed behaviors. Add a persona's first and last name (evocative, not stereotypical). Goals are the most critical detail, derived from analyzing the group of behaviors.
5. Check for completeness: Ensure the cast of personas covers the range of behaviors and user types identified in research.
6. Develop narratives: Third-person narrative is powerful for conveying attitudes, needs, and problems. A typical narrative should be one to two pages, not a short story. It quickly introduces the persona and briefly sketches a day in their life, including peeves, concerns, and interests relevant to the product. Choose photographs to make personas feel more real.
7. Designate persona types: There are six types, typically designated in this order:
- Primary: The primary target for the interface design. Only one per interface, though products may have multiple interfaces with different primary personas.
- Secondary: Satisfied by the primary persona's interface if one or two specific additional needs are addressed. An interface typically has zero to two secondary personas.
- Supplemental: User personas that are not primary or secondary, completely satisfied by the primary interface. Political personas often become supplemental.
- Customer: Address needs of customers, not end users. Typically treated as secondary personas, though may be primary for administrative interfaces.
- Served: Not users of the product but directly affected by its use. Track second-order social and physical ramifications. Treated like secondary personas.
- Negative: Not users of the product, used rhetorically to communicate who should not be the design target. Good candidates include technology-savvy early-adopter personas for consumer products and IT specialists for enterprise products.
⭐ Key Takeaways
The most critical concept is that personas are composite archetypes based on ethnographic research, not stereotypes or fictional characters, and they serve as the foundation for user-centered design by providing a precise, empathetic design target. Goals—specifically life, experience, and end goals—are the drivers of user behavior and must be inferred from qualitative data, not asked directly, making them essential for translating observed behaviors into design decisions. Personas resolve critical design problems including the elastic user, self-referential design, and designing for edge cases, while providing a common language for communication across the development team. The persona construction process follows seven systematic steps from revisiting hypotheses through designating persona types, with primary, secondary, and other types each serving specific roles in the design framework. Finally, personas must be context-specific, based on behavioral data (not demographics), and treated with dignity to avoid degrading into stereotypes.
🧠 Quick Revision Questions
- What are the three user-centered design issues that personas help resolve, and how does each issue typically manifest in product development?
- What are the three categories of user goals, and why must goals be inferred from qualitative data rather than asked directly?
- How do personas differ from user roles, user profiles, and market segments in terms of their basis and purpose in design?
- What are the seven steps in constructing personas, and what is the significance of the "clustering" step in Step 2?
- List and briefly describe the six types of personas, explaining the primary distinction between primary, secondary, and supplemental personas.