PSY504 — Midterm Summary (Lectures 1–22)
📘 Lecture 1 — INTRODUCTION
📖 Overview: This lecture introduces cognitive psychology as the scientific study of mental processes involved in thinking, acquiring, and storing knowledge. It traces the historical evolution of the field from ancient Greek philosophy through behaviorism to the modern information processing approach, explaining how cognitive psychology emerged as a distinct discipline.
🗂️ Topics Covered
The lecture begins by defining cognitive psychology and its focus on cognition as "thinking" or "knowing." It then explores the historical background starting with Plato's theory of knowledge acquisition, followed by the contributions of Wilhelm Wundt and early psychophysics. The lecture examines the limitations of introspective methods and the rise of behaviorism, before detailing the pivotal developments during and after World War II—including information theory, Noam Chomsky's critique of behaviorist language models, the emergence of computers and artificial intelligence, and Donald Broadbent's work on attention—that culminated in Ulric Neisser formally establishing cognitive psychology.
📝 Lecture Summary
Overview and Definition
Cognitive Psychology deals with cognition, which can be understood as "thinking" or "knowing." In other words, cognitive psychology deals with the processes involved in thinking, acquisition, and storage of knowledge. For this purpose, it adopts an information processing approach, viewing the mind as a system that processes information in stages similar to a computer.
Historical Background
Plato's Theory of Knowledge Plato, the great Greek philosopher, was the first person to present a coherent theory of how knowledge is acquired and retained. He proposed that ideas are created in the human mind and that these ideas are then projected out into the world. These projections serve as images that we see through our senses. In other words, the outside world is an illusion made up of projections of ideas and the true reality lies inside of us. Therefore, Plato concluded that perception is an internal process and we can learn everything by looking inwards.
Mental Philosophy and Epistemology When psychology was first taught in European universities, it was subsumed under the title of mental philosophy. Philosophers throughout history have been concerned with concepts of perception and knowledge, as to how these interact with reality. Similarly, the field of epistemology within philosophy has been concerned with the nature of knowledge. Thus cognitive psychology has been present as an undercurrent in the field of ontology and epistemology throughout the last two millennia.
Wilhelm Wundt and Psychophysics More recently, in 1875, Wilhelm Wundt set up the first psychological laboratory to study perception and cognition. A lot of the perceptual experiments and studies conducted were included in the field he called psychophysics. An example of psychophysics is the relationship of sensation and intensity of the stimulus.
🔑 Definition — Psychophysics: The study of the relationship between physical stimuli and the sensations and perceptions they produce.
📐 Formula: [No specific formula provided in text] → The lecture only provides the example: relationship of sensation and intensity of the stimulus.
Problems with Introspective Methods A major problem with most psychological studies of this time was over-reliance on introspective reports. In these reports, information was acquired by asking subjects what they felt, thought, or saw, heard, etc., and these reports were then used for deriving psychological principles. This was around the same time that Freud proposed the idea of unconscious processing. We now know for sure that most cognitive processing takes place at an unconscious level.
The Rise of Behaviorism The criticism of the introspective technique soon led psychology to its opposite extreme, and the behaviorist school took over. The behaviorists argued that anything we could not observe could not form part of the science of psychology. They coined and exclusively used terms like stimulus, response, reinforcement, conditioning, etc. All of these had to do with phenomena that could be converted into some kind of numerical representation. Hunger, for example, was called "number of hours of food deprivation." Behaviorists relied only on things they could see and rejected phenomena such as memory and imagery as unscientific just because they couldn't think of a way of observing and measuring them.
💡 Why this matters: Behaviorism's rejection of unobservable mental processes created a major gap in psychology, as it could not explain complex human behaviors like language, memory, and problem-solving.
The Cognitive Revolution During the 2nd World War, human factors research and information theory combined to generate the information processing approach. This approach, along with several other factors, led to the creation of a new field called cognitive psychology.
Among these factors was Noam Chomsky's critique of Skinner's book Verbal Behavior. Chomsky, in his groundbreaking paper "On Verbal Behavior," shattered the simple-minded behaviorist model of language designed by Skinner. He argued that language was far too complex to be explained by stimulus response alone.
Around the same time, computers had emerged as thinking machines, where a lot of similarities with human information processing were coming to the fore. The field of artificial intelligence had also emerged, which sought to make computers that thought like humans and solved problems and learned new things.
Donald Broadbent was working at the same time on attention and visual perception. A lot of experimental work during that time, along with Bartlett's classic experiments on memory, combined to create what Ulric Neisser called Cognitive Psychology in a book entitled Cognitive Psychology.
⭐ Key Takeaways
Cognitive psychology is the scientific study of thinking and knowing, adopting an information processing approach to understand how the mind acquires and stores knowledge. The field has deep philosophical roots extending back to Plato, but modern cognitive psychology emerged as a reaction against behaviorism, which had rejected the study of unobservable mental processes. The cognitive revolution was driven by multiple forces: Chomsky's critique proving behaviorism couldn't explain language, the development of computers and artificial intelligence as models for human thinking, and experimental work by researchers like Broadbent and Bartlett. The key insight is that most cognitive processing occurs unconsciously, and valid scientific methods can study internal mental processes through observable behaviors and information processing models. Students must remember that cognitive psychology is NOT behaviorism—it embraces the study of memory, imagery, attention, and other mental phenomena that behaviorists rejected as unscientific.
🧠 Quick Revision Questions
- What does the term "cognition" mean, and what approach does cognitive psychology adopt to study it?
- How did Plato explain the relationship between ideas, perception, and reality?
- What was the main limitation of introspective reports, and what did Freud propose around the same time?
- Why did behaviorists reject the study of memory and imagery, and how did Chomsky challenge the behaviorist model of language?
- Which three major developments during and after World War II contributed to the creation of cognitive psychology as a distinct field?
📘 Lecture 2 — The Information Processing Approach
📖 Overview: This lecture introduces the information processing approach, which focuses on how sensory input is transformed into behavioral output through mental processes. It explains why cognitive psychologists study what happens between sensation and behavior, using the computer analogy of hardware and software level descriptions to understand human cognition.
🗂️ Topics Covered
The lecture covers the information processing approach as an alternative to behaviorism, the concept of stages and levels of processing in the mind, the distinction between hardware level and software level descriptions of cognitive processes (with visual processing as an example), the role of attention and limited capacity models, and how cognitive psychologists develop and test theoretical models through experiments.
📝 Lecture Summary
The Information Processing Approach
Unlike behaviorism's stimulus-response model, the information processing approach examines how input is transformed into output. The key question is what happens between sensation and behavior. Cognitive psychology treats sensations as bits of information that undergo various mental processes, which may or may not result in behavior.
💡 Why this matters: This shift from behaviorism to information processing marks the foundation of modern cognitive psychology, allowing researchers to study internal mental processes scientifically.
Stages and Levels of Processing
These processes are usually performed in stages, with different layers or levels of processing in each stage. These layers can be understood as levels of description rather than actual processes themselves. Just as computers have hardware and software levels, human information processing can be described similarly:
- Hardware level description: What happens in the brain or nervous system when a sensation occurs
- Software level description: Mental processes like closing eyes to recall an image of a sensation
Hardware Level Description Example: Visual Processing
A hardware level description of visual processing includes:
- Studying the visual sensation itself
- How sensory neurons carry information higher in the system
- Studying the visual cortex (the brain part concerned with visual processing)
- Connections between visual cortex and other brain parts
- Complex brain processes we don't fully understand
- Afferent neurons carrying decisions via nerves to muscles that implement the decision
Software Level Description
A software level description would start at sensation and continue with:
- Sensory storage
- Possibility of a filter underlying selective attention
- Short term/working memory that processes and transforms material
- Storage in long term memory
- Retrieval from long term memory back into working memory when needed
Attention and Limited Capacity Models
The lecture uses the example of attention to show how doing two tasks simultaneously can impair performance quality at both tasks. This observation has allowed psychologists to generate limited resource or limited capacity models of attention.
Model Development and Testing
Cognitive psychologists generate descriptions of different stages of information processing and develop models that incorporate these descriptions into new theoretical frameworks. These models are then tested in the laboratory using experiments, mostly on human subjects.
⭐ Key Takeaways
The information processing approach fundamentally shifts focus from stimulus-response associations to understanding the internal mental transformations between input and output. Students must remember that cognition can be described at both hardware (neural/brain) and software (mental process) levels, similar to computer analysis. Attention operates as a limited resource system, meaning multitasking degrades performance. Cognitive models are built from these descriptions and tested experimentally on humans. Finally, the stages of processing—sensory storage, attention, working memory, and long term memory—form the backbone of cognitive psychology explored throughout this course.
🧠 Quick Revision Questions
- What is the fundamental difference between the information processing approach and behaviorism's stimulus-response model?
- What are the two levels of description used to understand human information processing, and what does each level describe?
- In the hardware level description of visual processing, what role do afferent neurons play?
- How does the example of doing two tasks simultaneously inform our understanding of attention?
- How do cognitive psychologists test the models they develop from information processing descriptions?
📘 Lecture 3 — Cognitive Neuropsychology
📖 Overview: This lecture introduces cognitive neuropsychology, which examines cognition at the hardware (neural) level of the brain. It explains how studying brain-injured humans, brains of deceased individuals, neuro-imaging, and animal studies contribute to our understanding of cognitive processes, and it details the fundamental biological units of the brain, including neurons and synapses.
🗂️ Topics Covered
Cognitive neuropsychology describes cognition at the hardware level, utilizing methods like studying brain-injured humans, brains of dead people, neuro-imaging (X-rays, MRI, fMRI), and animal studies. It then covers the structure and function of the neuron, including the cell body, dendrites, and axon, and explains the synapse and the role of neurotransmitters in transmitting impulses. Finally, it outlines the organization of the brain into its four main lobes: occipital, frontal, temporal, and parietal.
📝 Lecture Summary
Cognitive Neuro-psychology
Cognitive Neuro-psychology describes cognition at the hardware level, using the computer metaphor. The neural architecture of cognition is the basis on which the software level is built. At this level, it is possible to explain many visual and auditory phenomena, but higher-level cognitions remain a mystery. 💡 Why this matters: This lecture establishes that while we can observe brain activity patterns, the highest forms of thinking are not yet fully understood at a neural level.
Neuropsychological Methods: Brain-injured Humans
The study of brain-injured humans has greatly enriched our understanding of human cognition. It has allowed psychologists to design split brain experiments, which revealed the differences between the right and left hemispheres of the brain.
Neuropsychological Methods: Brains of Dead People
The study of brains of dead people has also added to the understanding of cognition, but to a limited extent. Researchers studied the brains of people with certain brain disorders to see if any traces of the illness could shed light on normal brain functioning.
Neuropsychological Methods: Neuro-imaging
X-rays have contributed to our understanding of brain processes, but Magnetic Resonance Imaging (MRI) and functional MRI (fMRI) scanning techniques have been even more revealing. The MRI technique is quite intrusive and yields relatively limited information. The fMRI, however, allows for live brain scans, is less intrusive, and is radiation-free. Despite these advances, it remains a hardware-level understanding of the brain and can never substitute a software-level description.
Neuropsychological Methods: Animal Studies
A controversial method of studying neural processes is studying live animals. This method is controversial because these animals are subjected to procedures like brain surgeries where parts of the brain are removed to observe their function. The conditions in which these animals are kept have also come into question. While a lot of useful information has been obtained, the ethical controversy remains.
The Neuron
The brain contains approximately 70 billion neurons. A neuron is a specialized cell that transmits and stores information of different kinds. The cell body contains a nucleus at its center, which governs the functions of the neuron. Tiny branches called dendrites bring information to the neuron from other neurons. On the other side, the neuron has a branch called the axon, which transmits information from the neuron to the muscles.
🔑 Definition — Neuron: A specialized cell that transmits and stores information of different kinds.
The Synapse
The neuron is not directly connected to other neurons. A fluid called the neurotransmitter moves between the dendrites of one neuron and the axon of another neuron. The gap between the neurons, which contains the neurotransmitter, is called the synapse. The synapse transmits the electric impulse generated in one neuron to the other neurons.
🔑 Definition — Synapse: The gap between neurons which contains the neurotransmitter, transmitting the electric impulse from one neuron to the next.
Organization of the Brain
The brain can be divided into four lobes: Occipital lobe, Frontal lobe, Temporal lobe, and Parietal lobe. Each lobe performs certain specialized functions.
⭐ Key Takeaways
The study of brain-injured humans, especially through split brain experiments, has been crucial for understanding hemispheric differences. While animal studies provide useful information, they remain ethically controversial. Neurons, composed of dendrites and axons, are the fundamental functional units, and they communicate across the synapse via neurotransmitters. The brain is organized into four main lobes, each with specialized functions. fMRI allows for live, non-invasive brain scanning, but provides hardware-level data, which cannot replace software-level descriptions of cognition.
🧠 Quick Revision Questions
- What is the core difference between hardware-level and software-level descriptions of cognition in this context?
- What are the four methods for studying brain processes discussed in this lecture?
- Name the three main parts of a neuron and their respective functions.
- What is the synapse and what is its role in neural communication?
- List the four lobes of the brain mentioned in the lecture.
📘 Lecture 04 — Cognitive Neuropsychology (Continued)
📖 Overview: This lecture continues the exploration of cognitive neuropsychology, focusing on the hardware level of cognition—the neural architecture underlying mental processes. It examines the structure and function of sensory systems, particularly vision and audition, and how they transmit information to the brain, providing a foundational understanding of how perceptual information is processed at the neural level.
🗂️ Topics Covered
The lecture covers the structure and function of the eye, including contributions by Ibn-al-Haitham, the visual pathway from retina to brain including the optic chiasma, lateral geniculate nucleus, and superior colliculus, the structure of the ear and the auditory pathway, and the neural architecture of sound localization through delay detectors.
📝 Lecture Summary
The Eye
Cognitive Neuropsychology describes cognition at the hardware level using the computer metaphor, where the neural architecture of cognition serves as the basis for the software level of mental processes. At this level, it is possible to explain many visual and auditory phenomena, though higher-level cognitions remain a mystery.
Visual information passes through the lens, which helps focus the image on the retina. From the retina, information travels to the optic nerve, which transmits it to the brain.
🔑 Definition — Hardware level: The neural architecture of cognition, analogous to computer hardware, that forms the physical basis for mental processes.
📐 Formula: Visual pathway → retina → optic nerve → brain (basic transmission sequence)
📌 Example: When light enters the eye, it passes through the lens which adjusts to focus the image precisely onto the retina, similar to a camera lens focusing light onto film.
💡 Why this matters: Understanding the hardware level helps explain why certain visual phenomena occur and provides the foundation for understanding visual disorders.
Ibn-al-Haitham
Ibn-al-Haitham, a great Muslim scientist, discovered not only the structure and function of the eye but also how it links with the nervous system. He proposed that the eyes transmit information to the brain via the optic nerve and was aware of different visual fields in the eye. He also proposed a dual visual pathway system. Among his other contributions were the development of spectacles and telescopes.
🔑 Definition — Dual visual pathway system: Ibn-al-Haitham's proposal that visual information follows two distinct neural pathways from the eyes to the brain, processing different features of the visual image.
📌 Example: Ibn-al-Haitham recognized that the left and right visual fields from each eye are processed separately before being combined in the brain, a finding that preceded modern understanding by centuries.
The Visual Pathway
The visual pathway starts from the retina where the image is formed, then travels to the optic nerve. From the optic nerve, information goes to the optic chiasma, where visual information from both eyes is combined. The combined information then moves to the lateral geniculate nucleus (LGN), which processes information about colors and details of the image. A separate visual pathway takes information about global features such as localization and movement to the superior colliculus. There is considerable evidence for two different visual pathways that take different features of an image to different parts of the brain.
🔑 Definition — Optic chiasma: The point where visual information from the two eyes is combined before being sent to the brain for further processing.
🔑 Definition — Lateral geniculate nucleus (LGN): A structure in the brain that processes information about colors and fine details of the visual image.
🔑 Definition — Superior colliculus: A brain structure that processes information about global visual features such as localization and movement.
📐 Formula: Visual pathway → retina → optic nerve → optic chiasma → lateral geniculate nucleus (colors/details) AND superior colliculus (localization/movement)
📌 Example: When viewing a moving red car, the lateral geniculate nucleus processes the car's red color and detailed features like the license plate, while the superior colliculus processes where the car is located and its direction of movement.
The Ear
The ear receives auditory information in the form of sound waves that impact the eardrum (tympanic membrane). From the eardrum, information is transmitted via the cochlea to the auditory nerve. The ear consists of three parts: the outer ear, middle ear, and inner ear, which includes the semicircular canals for balance and the cochlea for hearing.
🔑 Definition — Cochlea: The spiral-shaped structure in the inner ear that converts sound vibrations into neural signals sent to the brain via the auditory nerve.
📐 Formula: Auditory pathway → sound waves → eardrum → cochlea → auditory nerve → brain
📌 Example: When a bell rings, sound waves travel through the air, strike the eardrum causing it to vibrate, these vibrations are transmitted through the middle ear bones, then converted to neural signals in the cochlea before traveling to the brain via the auditory nerve.
The Auditory Pathway
The auditory pathway transmits auditory information from the ear to the brain for processing.
Sound Localization
Sound localization refers to the process of determining where a sound is coming from. This is enabled by delay detectors in the nervous system. These delay detectors are connected to each ear and process information about which ear received the sound first and which received it later. Depending on the delay between the ears, the system determines which direction the sound came from.
🔑 Definition — Delay detectors: Neural mechanisms that compare the timing of sound arrival at each ear to determine the direction of the sound source.
🔑 Definition — Sound localization: The neural process of determining the spatial origin of a sound based on the time difference between when it reaches each ear.
📐 Formula: Sound localization → time delay between ears → delay detectors → direction determination
📌 Example: If a person is standing to your left and calls your name, the sound reaches your left ear slightly before your right ear. The delay detectors in your brain detect this time difference and determine the sound came from the left side. If the delay is zero (sound reaches both ears simultaneously), the sound source is directly in front of or behind you.
💡 Why this matters: Sound localization is critical for survival (detecting threats) and daily functioning (identifying who is speaking in a crowd).
⭐ Key Takeaways
The lecture establishes that cognitive neuropsychology examines cognition at the hardware level, focusing on the neural architecture underlying mental processes. The visual system involves a complex pathway from the retina through the optic nerve and optic chiasma to either the lateral geniculate nucleus (processing colors and details) or the superior colliculus (processing movement and localization). Ibn-al-Haitham made foundational contributions to understanding the eye's structure and dual visual pathways, as well as inventing spectacles and telescopes. The auditory system processes sound waves through the eardrum and cochlea to the auditory nerve, with sound localization achieved through neural delay detectors that compare the arrival time of sound at each ear.
🧠 Quick Revision Questions
- What is the difference between what the lateral geniculate nucleus and the superior colliculus process in the visual pathway?
- What are the three main structures in the visual pathway from the retina to the brain, in order?
- How do delay detectors in the nervous system enable sound localization?
- What were Ibn-al-Haitham's key contributions to understanding the visual system?
- What happens to visual information at the optic chiasma?
📘 Lecture 5 — COGNITIVE PSYCHOLOGY (CONTINUED)
📖 Overview: This lecture explores how visual information is processed from the retina to the visual cortex, based on classical experiments by Kuffler, Hubel & Wiesel, and David Marr. It then introduces sensory memory, focusing on iconic (visual) memory and Sperling's groundbreaking partial-report procedure, which revealed the capacity and rapid decay of visual sensory storage.
🗂️ Topics Covered
The lecture covers information processing in ganglion cells (on-off and off-on cells), Hubel & Wiesel's discovery of bar and edge detectors in the visual cortex, David Marr's computer model of visual processing, an introduction to sensory memory (including the five senses), iconic memory and Sperling's whole-report experiment, Sperling's partial-report procedure demonstrating the high capacity of visual sensory memory, and the decay of visual memory over time.
📝 Lecture Summary
Information Processing in Visual Cells
Information processing of visual information occurs in our visual cortex. Classical experiments were conducted to understand how sensation transforms into perception. In 1953, Kuffler studied ganglion cells and discovered on-off cells and off-on cells. These cells define visual information processing. If light fell on the centre of the retina, on-off cells were activated; if light fell on the periphery, off-on cells started firing. When light is on the centre, on-off cells work and off-on cells do not. When there is illumination on the surrounding point, on-off cells become silent and off-on cells start working. When there is diffuse illumination, both cells fire slowly. This processing happens at the ganglion cell level, before information reaches the visual cortex.
Hubel & Wiesel
In 1962, Hubel & Wiesel conducted an important experiment using cats. They showed different stimuli, lights, and shapes to cats and viewed the visual cortex. They found that visual cortical cells respond in a more complex manner than lower cells. They named these cells bar detectors and edge detectors. Edge detectors help us understand where an object ends and another starts. They respond positively to light on one side of a line and negatively to light on the other side. Bar detectors respond positively to light in the centre and negatively to light at the periphery, or vice versa. Both bar and edge detectors are specific with respect to position, orientation, and width — they respond only to stimulation in a small area of the visual field. Bar and edge detectors combine to help us see many objects.
David Marr’s Work
David Marr developed a computer model of how information from on-off and off-on cells could be used to yield a useful analysis of the visual image. Marr and Hildreth (1980) combined the output of on-off and off-on detectors to calculate bars and edges of various widths and orientations. In this computer system, symbolic descriptions were created. The boundaries of objects in real images pose a difficult problem in computer vision. This system combines symbolic descriptions to identify the contour of an object.
Sensory Memory
When information first enters the human system, it is registered in sensory memories. Sensory memory allows us to take a snapshot of our environment and store this information for a short period. Only information transferred to other levels of memory will be preserved, and it lasts no more than about two seconds. Sensory memory holds a short impression of sensory information even after the sensory system stops sending information. There are 5 basic senses: 1) Vision, 2) Hearing, 3) Smell, 4) Taste, and 5) Touch. Most research has focused on vision and hearing because technology is advanced in these areas (e.g., cameras, microphones, computers).
Visual Sensory Memory — Iconic Memory
Iconic memory (visual sensory memory) can store a great deal of information but only for a very brief period. In an experiment, a subject sits looking at a screen. A dot is shown in the centre, and the subject focuses on it. A set of letters is presented at that fixed point for a very brief period, then removed. Subjects are asked to report the items. Normally, subjects report only 4–5 items out of 12 letters.
🔑 Definition — Iconic Memory: A brief visual sensory store that can effectively hold all the information in a visual display for a very short period (less than one second).
Sperling’s Partial Report Procedure
Sperling conducted an experiment in visual sensory memory using the same array of letters but asked subjects to report letters according to a cue (an auditory beep). After the array was turned off, a tone was sounded: high tone cue for reporting the top row, medium tone for the middle row, and low tone for the bottom row. Subjects reported at least 3 out of 4 letters, no matter which row was cued. Since subjects did not know beforehand which row would be cued, they had at least 3 letters from each row available — meaning they had at least 9 out of 12 letters available in their visual memory. This method was called the partial-report procedure.
💡 Why this matters: This experiment demonstrated that the capacity of iconic memory is much larger than previously thought — we briefly hold nearly all visual information, but it decays too quickly to report it all in whole-report tasks.
The Decay in Visual Memory
Sperling also varied the length of the delay between the offset of the display and the tone. As the delay increased to 1 second, subjects' performance decayed back to the original whole-report level of 4 or 5 items. This indicates our visual memory loses most of the information in one second. A criticism of this experiment is that it is an artificial or laboratory phenomenon, not related to real-life vision. Sperling's experiments indicate the existence of a brief visual sensory store — a memory that can effectively hold all the information in the visual display. We receive a lot of visual information from our surroundings, but it is all stored in visual sensory memory for only a very brief period. Information that receives attention is stored in other memory systems; the rest is lost.
⭐ Key Takeaways
Kuffler's discovery of on-off and off-on ganglion cells and Hubel & Wiesel's identification of bar and edge detectors in the visual cortex demonstrate how the brain processes basic visual features hierarchically. David Marr's computer model showed how these neural outputs can be combined to recognize object boundaries. Sensory memory, particularly iconic (visual) memory, briefly holds a high-capacity snapshot of our environment. Sperling's partial-report procedure revealed that iconic memory can store nearly all visual information (e.g., 9 out of 12 letters) but decays rapidly — within about one second — returning to whole-report levels of only 4–5 items.
🧠 Quick Revision Questions
- What are on-off cells and off-on cells, and under what lighting conditions does each fire?
- How do edge detectors and bar detectors differ in their response patterns?
- What was the key finding of Sperling's partial-report procedure compared to the whole-report procedure?
- How does the delay between stimulus offset and the cue affect performance in Sperling's experiment, and what does this reveal about iconic memory?
- What are the five basic senses, and why has most sensory memory research focused on vision and hearing?
📘 Lecture 06 — VISUAL SENSORY MEMORY EXPERIMENTS (CONTINUED)
📖 Overview: This lecture continues the exploration of visual sensory memory through Sperling’s classic experiments, revealing how post-exposure field brightness affects memory duration. It introduces the concept of iconic memory and echoic memory, and presents fascinating evidence that our perception of the present is actually delayed — a phenomenon known as psychological time.
🗂️ Topics Covered
This lecture covers Sperling's 1967 variation using dark versus light post-exposure fields and the resulting increase in retention time to 5 seconds. Neisser's coining of the term iconic memory for brief visual storage is explained, along with the overwriting/erasure effect. The lecture then introduces psychological time — the idea that we perceive the present after a delay. Finally, auditory sensory memory (echoic memory) is examined through the experiments of Moray, Bates & Barnett (1965) and Darwin, Turvey & Crowder (1972).
📝 Lecture Summary
Sperling (1967) & Neisser (1967)
Sperling made another variation of his classic experiment: after the array of letters disappeared, he made the visual field dark instead of white. This produced fascinating results — the retention power of the subjects increased to 5 seconds. When the post-exposure field was light, sensory information remained for only 1 second, but when the field was dark, it remained for a full 5 seconds.
Ulric Neisser wrote the first cognitive psychology book (1967). He coined the term word icon and devised the term iconic memory for short-term visual memory. According to Neisser, visual memory is neither short-term memory nor long-term memory — it is a very, very brief memory and should be called iconic memory. Without such a visual icon, perception would be much more difficult. Many stimuli are of very brief duration; to recognize them, the system needs some means of holding onto them for a short while until they can be analyzed.
Neisser also reported that if another display is presented during that one second when you are already retaining an image, it is like erasing or washing out the first and overwriting it with the second. Almost all information is held for a very brief period (1 second). It is quickly washed out after removal of the stimulus unless attention is paid to it.
The sensory store is particularly visual in character and is sensitive to light. We cannot pinpoint where sensory memory is located in the brain — this is a software-level description, not a hardware-level description. The sensory visual store is not physical but seems like a physical phenomenon and is sensitive to light.
💡 Why this matters: This distinction between light and dark post-exposure fields reveals that iconic memory is fundamentally a sensory-level phenomenon — it decays when light overwrites it, but persists in darkness. This explains why afterimages seem to linger longer in dark rooms.
Psychological Time
Are we living in every moment in this moment, or are we living in the past? The conclusion drawn is that our new information synthesizes with our old information as it happens at the same time and space. If we see things that happened one second ago, we perceive them as happening here and now. We are attending to visual information after some delay — no matter how brief that period may be (less than one second). And because we perceive the information after we have seen something, we know we are not living in the past; we see things here and now.
This is called psychological time. We found that our perception is delayed by about a second. Our visual system is recombining things within a second and putting them together, constructing a visual image based on the information we receive.
Auditory Sensory Memory
Evidence for an auditory sensory memory similar to visual memory comes from experiments by Moray, Bates and Barnett (1965) and Darwin, Turvey & Crowder (1972).
Moray, Bates and Barnett (1965)
Their experiment is a copy of Sperling's experiments, but instead of top, middle, and bottom, sounds were manipulated as coming from left, right, or front. The cue was presented visually (the opposite of Sperling's experiment, where the cue was auditory). In their experiment, subjects listened to a recording over stereo headphones, hearing three lists of three items read simultaneously. Because of stereophonic mixing, one list seemed to come from the left side of the subject's head, one from the middle, and one from the right side.
The investigators compared results from a whole-report procedure (subjects reported all nine items) with a partial-report procedure (subjects were cued visually after presentation to report items from left, middle, or right). A greater percentage of the letters were reported in the partial-report procedure than in the whole-report procedure. This was statistically significant — an obvious, marked difference.
Thus, all the information is available to short-term sensory storage, but it quickly decays. The delay is significant but not as striking as visual sensation because visual sensation is central and hearing only supports it. Neisser (1967) has called this auditory memory echoic memory.
Neisser pointed out many things we understand and perceive but cannot perform — like the difference between competence and performance. This was a big blow to behaviorism. Cognitive psychology scientifically proved that psychology is much more than observable behavior only. Neisser's conclusion: the whole-report paradigm was a limit to our production, not a limit to our process. We process everything but we cannot produce all. Time delay, overwriting, and light contrast intervene in the production aspect, not in the storage aspect.
💡 Why this matters: The whole-report vs. partial-report distinction reveals that our sensory systems register far more information than we can consciously report. The bottleneck is in production (reporting), not in processing.
Psychological Time (Recap)
We all are living in the past. This past is not in passive terms. We receive visual information that is one second old, and we receive auditory information that is 5 seconds old. Information is not only being received but is also being reconstructed. We are creating reality afresh every moment. Every icon is a separate entity. New information is a different entity. It reconstructs reality by combining with other entities. The heart of cognitive psychology lies in its experimentation.
⭐ Key Takeaways
The most critical point from this lecture is that iconic memory persists for ~1 second in a light post-exposure field but up to ~5 seconds in a dark field, and echoic memory (auditory sensory memory) persists even longer. Neisser's distinction between competence (what we process) and performance (what we can report) revolutionized cognitive psychology by challenging behaviorism — showing that the whole-report paradigm measures a production limit, not a processing limit. The concept of psychological time reveals that our perception of "now" is actually a reconstruction of information ~1 second old (visual) and ~5 seconds old (auditory), meaning we are constantly creating reality afresh from delayed sensory input. Finally, sensory memory is a software-level description, not a hardware location — it is sensitive to light and sound but cannot be localized to a specific brain area.
🧠 Quick Revision Questions
- What was the difference in retention time for iconic memory between a light post-exposure field and a dark post-exposure field in Sperling's variation?
- According to Neisser, what distinguishes iconic memory from short-term memory and long-term memory?
- What does the concept of psychological time reveal about how "present" our perception actually is?
- In Moray, Bates, and Barnett's (1965) experiment, how was the procedure different from Sperling's original experiment, and what was the key finding?
- What did Neisser mean when he said the whole-report paradigm was "a limit to our production, not to our process"?
📘 Lecture 07 — ATTENTION
📖 Overview: This lecture introduces attention as a fundamental cognitive process, exploring its limited-capacity nature through various metaphors and experimental evidence. It examines how attention selects information from sensory memory, the limitations of this selection process, and classic experimental paradigms like dichotic listening that reveal how attention operates, including debates about whether selection occurs early or late in processing.
🗂️ Topics Covered
The lecture covers the conceptualization of attention as a limited mental resource using multiple metaphors (spotlight, filter, bottleneck, energy, workspace, demons). It discusses the single-minded nature of attention and its capacity limitations, explores how these limitations manifest in sensory tasks like whole report performance, presents the dichotic listening paradigm by Cherry and Moray, introduces the Early Selection Model of attention, and examines research by Gray and Wedderburn showing that attention can track meaning across ears.
📝 Lecture Summary
What is attention?
Attention is conceived of as being a very limited mental resource. It plays a crucial role in the stage of selection of information from sensory memory, where most information is discarded and only some is selected for further processing. Numerous metaphors help us think about the limited-resource characteristics of attention.
Common metaphors for attention include:
- Spotlight/beam – like a stage spotlight that falls on one person, attention focuses on one input
- Filter/sieve – like listening to one conversation at a party while filtering out others
- Bottleneck/narrow lane – like traffic that can only flow through a narrow passage
- Limited resource – like skilled manpower or petrol, only enough for limited tasks
Additional limited-resource metaphors:
- Energy: Like electricity with a fixed current, attention has a fixed energy supply allocable to only so many tasks. If allocated to more, performance degrades or a "fuse would blow."
- Spatial or workspace: Like a small room that can only fit a few employees, limited workspace means only so many tasks can be performed.
- Demons: Attention as a small set of agents (demons) that can perform tasks but only one at a time.
🔑 Definition — Attention: a very limited mental resource that selects information from sensory memory for further processing.
Single-mindedness
Attention is single-minded in sense, not double-minded. This single-mindedness means that only enough energy, only enough workspace, or only a single attention demon is available for one task or process. So it means one thing at a time. Attention has a capacity to perform only one demanding task. Attention cannot perform two demanding tasks simultaneously.
But what about walking and talking? Driving and talking, and smoking? Tasks that are practiced to the point at which they do not make excessive demands can be performed simultaneously. We cannot simultaneously do mental addition and carry on a conversation because each activity involves multiple attention-demanding subcomponents. Whether or not attention is truly single-minded, its capacity is severely limited.
📌 Example: Walking and talking can be done together because walking is so practiced it makes few demands, but mental addition and conversation cannot be done simultaneously because both involve multiple attention-demanding subcomponents.
💡 Why this matters: Understanding the single-minded nature of attention explains why multitasking often fails and why practice can make some tasks automatic, freeing up attentional resources.
Limitations in sensory tasks
The limited capacity of attention is the root cause of the reporting limitations demonstrated in visual and auditory reporting tasks. Attention appears to be the real reason for whole report performance. Limitations of the icon and echo are actually the limitations of attention. All the information gets into sensory memory, but to be retained, each unit of information must be attended to and transformed into some more permanent form. Sensory information needs to be stored in a form that can be retained. Selection for this transformation happens only to attended items. Most information is lost because attention has limited capacity.
Dichotic Listening Tasks
Cherry (1953) and Moray (1959) conducted experiments on how subjects select what sensory input they attend to, using a dichotic listening task.
In a dichotic listening experiment, subjects wear a set of headphones. They hear two messages — one ear presented with one message, the other ear with another message. Subjects pay attention to (shadow) one message and tune out the other. Psychologists discovered that very little about the unattended message is processed in a shadowing task. Subjects cannot tell what language was spoken or report any of the words spoken even if the same word was repeated over and over again. Subjects reported hearing very little from the other ear.
📌 Example (Shadowing Paradigm):
- Messages in Left Ear: ran, house, Ox, Cat
- Messages in Right Ear: Tea, job, books, look
- People asked to attend to left ear report: ran, house, Ox, Cat
- They report nothing from the unattended ear
Early Selection Model
Early selection means information is selected early, soon after the sensory store. It happens early on in this system. Higher level processing is not done. The meanings of the words do not have an impact at this stage in this model. Attention itself is a lower level process.
The Early Selection Model shows information coming from right and left ears, but before going to bottom line, information from one ear blocks the other ear's information. According to this model, attention is an early selection process and a low-level process.
🔑 Definition — Early Selection Model: a model proposing that attention selects information for processing soon after the sensory store, before higher-level meaning analysis occurs.
Attention and meaning
Gray and Wedderburn (1960) conducted an experiment at Oxford demonstrating that subjects were quite successful in following a message that jumped back and forth between ears. This challenged the early selection view.
📌 Example (Gray & Wedderburn's experiment):
- Left Ear: John — Eleven — books
- Right Ear: Eight — writes — Twenty
- Instructed to shadow the meaningful message, subjects reported: John writes books
Thus, subjects are capable of shadowing a message on the basis of meaning rather than physical ear. This suggests that some semantic processing of unattended information may occur before selection.
⭐ Key Takeaways
Attention is a severely limited mental resource that selects information from sensory memory for further processing, as illustrated by multiple metaphors including spotlight, filter, bottleneck, energy, workspace, and demons. While attention is single-minded and cannot perform two demanding tasks simultaneously, practiced tasks can become automatic and require fewer attentional resources. The limited capacity of attention explains why most sensory information is lost — each unit must be attended to for retention. Dichotic listening experiments (Cherry, Moray) show that very little of unattended messages is processed, supporting early selection views. However, Gray and Wedderburn's findings that subjects can shadow messages based on meaning across ears suggest that at least some semantic processing may occur before selection, challenging simple early selection models.
🧠 Quick Revision Questions
- What are the five main metaphors used to describe attention's limited-resource characteristics, and what does each emphasize about attention?
- Why can people walk and talk simultaneously but cannot do mental addition and carry on a conversation at the same time?
- In Cherry and Moray's dichotic listening experiment, what information were subjects able to report from the unattended ear?
- According to the Early Selection Model, at what point in processing does attention select information, and what level of processing occurs before selection?
- How did Gray and Wedderburn's 1960 experiment challenge the Early Selection Model of attention?
📘 Lecture 8 — Attention and meaning
📖 Overview: This lecture continues the discussion of attention by examining whether attention operates as an early or late selection process. Groundbreaking experiments by Gray & Wedderburn and Treisman demonstrated that meaning, not just physical characteristics, can guide selective attention, leading to the development of competing theoretical models.
🗂️ Topics Covered
This lecture covers the Gray & Wedderburn experiment showing attention can use meaning, Treisman's experiment on early versus late selection, three major attention models including Treisman's attenuator model, Broadbent's filter model, and Norman's late selection model, and concludes with a comparison of early and late selection theories.
📝 Lecture Summary
ATTENTION (continued)
Two undergraduates at Oxford, Gray and Wedderburn (1960) conducted an experiment challenging the idea that attention is a low-level early selection process. If attention selects only on physical characteristics like ear of presentation, then meaning should not matter. Subjects were given different words to each ear simultaneously: • Left Ear: John Eleven books • Right Ear: Eight writes Twenty
When instructed to shadow the meaningful message, subjects reported: "John writes books." This demonstrated that subjects are capable of shadowing a message based on meaning rather than physical ear, switching between ears to follow the meaningful content.
Implications
The Gray and Wedderburn experiment proved that attention can use meaning, which is a higher-level process. Subjects switched some information and selected meaningful words regardless of whether they were presented in the left ear or right ear. This suggests that attention is probably not an early selector but a late selector that operates after meaning has been processed.
Triesman's experiment
Treisman (1960) conducted an experiment to determine whether attention is an early selector or late selector. Subjects were asked to shadow the message in the left ear. The messages were: • In Left ear: I was going there when/ China, smoke, lovely, chirping • In Right ear: books, chairs, tables, elephant/ I saw a bright flash
Many subjects switched ears to follow the meaningful message, replicating the finding that meaning guides attention.
Attention models
1. Treisman's Model
Treisman proposed an early selection model based on her experiment. In this model, there is no filter but an attenuator. The attenuator weakens unattended signals rather than completely blocking them. After the attenuator, information is processed in a mental dictionary where some items are more important than others (e.g., our name is more important than "table" or "fire"). This model explains selection on the basis of meaningful information.
🔑 Definition — Attenuator: A device in Treisman's model that weakens unattended signals instead of completely filtering them out. 💡 Why this matters: The attenuator explains how we can still notice important information (like our name) from unattended channels, whereas a strict filter would block it completely.
2. Broadbent's experiment
Broadbent's model shows that senses take in information within their limited capacity. This information goes into short-term store (sensory store), which has unlimited capacity but retains information for only a very brief period. Then a selective filter tunes out extra information and passes selected information through a limited capacity channel to other systems. One path goes back to short-term store as a feedback loop. The limited-capacity channel also passes information to the store of conditional probabilities of past events, which then moves again towards the selective filter. Information then goes to a system for varying output until some input is secured, and then moves towards effectors (physical responses to attention).
🔑 Definition — Selective Filter: In Broadbent's model, a mechanism located soon after the short-term store that selects information based on physical characteristics, not meaning. 📐 Key feature: Broadbent's model does NOT select on the basis of meanings — selection occurs early, before semantic processing. 💡 Why this matters: This is the original "early selection" model and served as the foundation for all subsequent attention theories.
3. Norman's Model
Norman proposed a late selection model. He said we get all information that is stored in short-term memory, then in long-term memory. In long-term memory, we check its relevance; when relevance is determined, we select information and then pay attention. This model is opposite to early selection models because all information is processed for meaning before selection occurs.
🔑 Definition — Late Selection Model: A model where all incoming information is processed for meaning before attention selects what to focus on.
There are two most important models:
1. Early selection model — comes in two types:
- Filter model (Broadbent): Some things pass, some are stopped completely
- Attenuator model (Treisman): Some information is weakened, some is strengthened
2. Late selection model
In the attenuator model, the mental dictionary contains items that are very important to us. These important items are strengthened by the attenuator, allowing them to capture attention even from unattended channels.
🔑 Definition — Filter vs. Attenuator: A filter completely blocks unattended information; an attenuator weakens it so important information can still be noticed. 📐 Key distinction: Early selection occurs before meaning analysis; late selection occurs after meaning analysis.
⭐ Key Takeaways
The most critical thing to remember is that attention can operate based on meaning, not just physical characteristics, as proven by Gray & Wedderburn's and Treisman's experiments where subjects switched ears to follow meaningful messages. The two major competing theories are early selection (Broadbent's filter model and Treisman's attenuator model) and late selection (Norman's model). Treisman's attenuator model is particularly important because it explains how meaningful information like one's own name can break through from unattended channels. Broadbent's original filter model places selection early, before semantic processing, while Norman places it after full semantic analysis in long-term memory.
🧠 Quick Revision Questions
- What did Gray and Wedderburn's "John Eleven books / Eight writes Twenty" experiment demonstrate about attention?
- What is the key difference between Broadbent's filter model and Treisman's attenuator model?
- In Treisman's model, what is the role of the "mental dictionary" in selective attention?
- What is the fundamental difference between early selection and late selection models of attention?
- According to Norman's model, where does the selection of information for attention occur?
📘 Lecture 09 — Capacity Models
📖 Overview: This lecture examines attention through the lens of capacity models, shifting from earlier bottleneck theories to more flexible frameworks. It explores Kahneman's capacity model, which proposes a general limit on mental work capacity that can be flexibly allocated, and introduces the multimode theory that integrates capacity with stage of selection. Understanding these models is crucial for explaining how people manage multiple tasks and why interference occurs.
🗂️ Topics Covered
This lecture covers the shift from bottleneck to capacity models of attention, Kahneman's capacity model including arousal, enduring disposition, and momentary intentions as allocation policies, the World Trade Centre example illustrating attention allocation failure, comparison between bottleneck and capacity theories, the interaction between capacity and stage of selection, and the multimode theory which proposes that selection can occur at either early or late stages depending on task demands and intentions.
📝 Lecture Summary
Attention (continued)
Psychologists have become more interested in the capacity demands of different tasks (Kahneman, 1973). Different tasks require different amounts of attention, and attention can be diverted from one task to another. Tasks require mental effort. People may have some control over where the bottleneck occurs (Johnston & Heinz, 1978). 💡 Why this matters: This introduces the idea that attention is not fixed but can be flexibly managed, challenging earlier rigid filter models.
Kahneman’s Capacity Model
"Attention and Effort" was a major work of Kahneman (1973). He shifted the focus from bottleneck to capacity. There is flexibility in attention — we can change our attention from one thing to another. Evidence shows our bottleneck is actually adjustable and can move from early to late. There is a general limit on a person’s capacity to perform mental work. A person has considerable control over how this capacity is allocated.
🔑 Definition — Daniel Kahneman: A pioneer of cognitive psychology who has been working at Princeton University, contributing not only to attention but other areas.
Kahneman’s Capacity Model (Diagram Explanation)
In this model, many miscellaneous determinants impact the sensory system. Something happens that triggers arousal — meaning some activity starts. Arousal has many manifestations. Available capacity of attention will be allocated depending on the state of arousal. There is an allocation policy. Then there are some possible activities — the amount of attention paid on a task. There is a feedback loop: we evaluate how much attention is needed for the task, then readjust more capacity to that task.
🔑 Definition — Arousal: A physiological state that influences the distribution of mental capacity in the various tasks.
🔑 Definition — Enduring Disposition: An automatic influence where people direct their attention.
🔑 Definition — Momentary Intentions: A conscious decision to allocate attention to certain tasks or aspects of the environment.
📌 Example: World Trade Centre A Boeing 707 was flying in cloudy weather at an altitude 200 feet below the top of the WTC. The airport controller reported this. An alarm buzzed in the airport control tower to signal danger. The controller radioed the crew to turn around and climb up to 3000 feet. The controller was monitoring seven other planes at the time. His attention was diverted to other planes, so he could not pay attention to that plane. The alarm diverted his attention toward that plane.
Bottleneck vs. Capacity
Both models predict that simultaneous activities are likely to interfere with each other, but they attribute the interference to different causes.
- Bottleneck: According to this model, the same mechanism is required to perform 2 incompatible tasks. It is specific — it says the same mechanism is needed for those tasks.
- Capacity: Demands of 2 activities exceed available capacity, then there is a problem. It is non-specific to the task. Total demands of the task is an important variable.
Both kinds of interference occur (specific and capacity, early selection and late selection). Both kinds of theories are necessary.
Capacity & Stage of Selection
All experiments show flexibility of attention — we can divert our attention from one task to another. These also show an interaction between bottleneck and capacity theories (Johnston & Heinz, 1978). The listener has control over the location of the bottleneck. The location can vary from early mode of selection (before recognition) to late mode of selection (after semantic analysis, meaning meaning analysis). There is a need to combine capacity model and stage of selection.
Multimode Theory
A theory that proposes that people’s intentions and the demands of the task determine the information processing stage at which information is selected. Demands of the task require greater mental efforts. People’s intentions and the demands of task decide where attention is paid. We select on the basis of meanings if there are two competing tasks.
Early Selection vs. Late Selection
According to this theory, both early and late selection can occur. Attention is flexible — we can move our bottleneck and filter. But there must be interaction between bottleneck and capacity. We can shift our bottleneck to the low level in information processing when we pay attention to the physical properties of the task. We have a capacity to switch our bottleneck from early to late or late to early. Late selection will affect the perception of primary message because more information is selected about the secondary task.
⭐ Key Takeaways
Kahneman's capacity model revolutionized attention theory by shifting focus from a fixed bottleneck to a flexible capacity system with a general limit on mental work. The allocation of attentional capacity depends on arousal, enduring dispositions (automatic), and momentary intentions (conscious decisions). Bottleneck and capacity theories are both necessary: bottleneck explains specific interference from incompatible tasks requiring the same mechanism, while capacity explains non-specific interference when total demands exceed available capacity. The multimode theory (Johnston & Heinz, 1978) integrates these views, showing that people can control the location of the bottleneck — shifting between early selection (physical properties) and late selection (semantic analysis) depending on task demands and intentions. Late selection consumes more capacity because more information is processed about secondary tasks, which can impair primary task perception.
🧠 Quick Revision Questions
- What did Kahneman (1973) shift the focus of attention research from, and what did he propose instead?
- What are the three factors in Kahneman's model that influence the allocation policy of attention?
- In the World Trade Centre example, why did the controller fail to notice the Boeing 707's dangerous altitude?
- According to the multimode theory, what determines whether selection occurs at an early or late stage?
- How does late selection affect the perception of the primary message compared to early selection?
📘 Lecture 10 — Multimode Theory (continued)
📖 Overview: This lecture continues the discussion of Multimode Theory, which proposes that people’s intentions and task demands determine the stage at which information is selected for processing. It presents a key experiment testing this theory using selective listening with a subsidiary reaction-time task, and examines how capacity demands change based on whether selection occurs early (physical cues like pitch) or late (semantic cues like meaning).
🗂️ Topics Covered
This lecture covers the continuation of Multimode Theory, the main assumption of Capacity Theory, the concept of reaction time and accuracy, a detailed selective listening experiment with conditions of no list, one list, and two lists (using pitch or meaning), the predictions of Multimode Theory, results for both subsidiary and primary tasks, and the implications and summary of findings. The experiment demonstrates that selective attention requires capacity, that more capacity is needed for late selection (meaning) than early selection (pitch), and that reaction times and error rates increase accordingly.
📝 Lecture Summary
Multimode Theory (continued)
A theory which proposes that people’s intentions and the demands of the tasks determine the information processing stage on which information is selected. This means that selection can occur early, based on physical features, or late, based on semantic meaning, depending on what the task requires.
Experiments to test multimode theory A series of 5 experiments was conducted to measure the amount of capacity required to perform a task and to record how quickly a person responds to a subsidiary task (reaction time). The main task was a selective listening task, where different voices are presented in both ears, and subjects are asked to attend to only one voice or list and repeat it. The subsidiary task was to respond to a randomly appearing light signal by pressing a button.
🔑 Definition — Subsidiary Task: A task that typically measures how quickly people can react to a target stimulus in order to evaluate the capacity demands of the primary task.
Every theory has assumptions; if an experiment fulfills those assumptions, then the theory is supported. Every theory predicts some phenomenon, and these predictions are tested through experiments. If predictions are met, the theory is accepted; otherwise, it is rejected.
Capacity Theory
The main assumption of Capacity Theory is: The greater the portion of capacity allocated to selective listening (the primary task), the less should be available for monitoring the signal light, causing longer reaction times. In other words, if the selective listening task requires more capacity, then the subsidiary task's capacity to monitor the light signal is reduced, leading to a longer reaction time.
💡 Why this matters: This assumption directly links attention to a limited pool of cognitive resources, implying that tasks compete for a finite capacity.
Reaction Time & Accuracy
The time it takes for a subject to respond to a stimulus is called reaction time. The longer the time taken, the more difficult the task. If the task is demanding, subjects are likely to make mistakes and take a long time. With a more demanding task, there are more mistakes and less accuracy. A less demanding task leads to more accuracy. Sometimes reaction time is slow but accuracy is still present.
Experiment
In this experiment, pairs of words were presented simultaneously to both ears through headphones. Undergraduates at the University of Utah were asked to attend to different voices. The stimuli were selected either by the pitch of the voice (male/female) or by semantic category (cities, occupations).
The conditions of the experiment were:
- No lists: No list of stimuli was presented in either ear.
- One list: Only one list was presented in one ear (e.g., a male or female voice).
- Two lists: Two different lists of stimuli were presented in both ears. In this condition, subjects had to first understand meanings (late selection). In one ear, a male or female voice was presented; in the other ear, different city names or occupation names were presented.
The rationale was: those using pitch (male/female) were using physical information and could use an early, sensory mode of selection because the two messages were physically different. Those using meaning (cities, occupations) had to use a late, semantic mode of selection because it was necessary to know the meaning of the word before categorizing it.
Predictions of Multimode theory
Multimode theory predicts that more capacity is required to perform at a late mode of selection. Therefore, use of the semantic mode would cause slower reaction times to the light signal and more errors on the selective listening task. Late selection requires more capacity. In the no list condition, performance is best; in the one list condition, performance is better; in the two lists condition, performance is worst.
Results of subsidiary tasks
Performance on Subsidiary Task
- When No list was given, the reaction time was 310 ms
- When there was one List (either female voice or cities names), the reaction time was 370 ms
- In Two Lists condition (pitch), the reaction time was 433 ms
- In Two Lists condition (meaning), the reaction time was 482 ms
Results of Primary Task
Percentages of Errors
- In case of one List, the error percentage was 1.4%
- In case of Two Lists (Pitch), the error percentage was 5.3%
- In case of Two Lists (Meaning), the error percentage was 20.5%
Implications
The implications of this experiment are:
- Selective Attention requires capacity — it is not cost-free.
- Reaction Time is slower for two lists condition over one — more information to process increases difficulty.
- Amount of capacity required increases from early to late selection — semantic processing consumes more resources than physical feature processing.
- Reaction Time is slower for meaning than for pitch — categorizing by meaning is more demanding than differentiating by voice.
Summary
- Attention is flexible — people have the choice of how best to use it. People can allocate attention to a task according to their will.
- Task difficulty decreases with practice — over time, tasks become less demanding and require less capacity.
⭐ Key Takeaways
This lecture demonstrates that attention is not a fixed filter but a flexible, capacity-limited system. The core finding is that the demands of the primary task determine how much capacity remains for secondary tasks, with late-selection (semantic) tasks requiring more capacity than early-selection (physical) tasks. The experiment clearly shows that reaction times and error rates increase systematically as the complexity of the selective listening task increases, from no list (fastest, fewest errors) to two lists with meaning (slowest, most errors). Students must remember that Capacity Theory assumes a finite pool of attentional resources, and that Multimode Theory predicts that late selection consumes more capacity than early selection. The specific numerical results (310 ms, 370 ms, 433 ms, 482 ms for reaction times; 1.4%, 5.3%, 20.5% for errors) are essential for understanding the pattern of evidence supporting these theories.
🧠 Quick Revision Questions
- What is the main assumption of Capacity Theory regarding the relationship between the primary task and the subsidiary task?
- In the experiment described, what were the two different modes of selection (early vs. late) and what cues did they rely on?
- What were the reaction times for the subsidiary task under the "no list," "one list," "two lists (pitch)," and "two lists (meaning)" conditions?
- Why does Multimode Theory predict that the semantic (meaning) mode requires more capacity than the pitch mode?
- What were the error percentages for the primary task under the "one list," "two lists (pitch)," and "two lists (meaning)" conditions, and what does this suggest about task difficulty?
📘 Lecture 11 — RECAP OF LAST LESSONS
📖 Overview: This lecture revisits bottleneck and capacity models of attention, then introduces automaticity—how well-practiced skills require minimal mental effort. It examines the pros and cons of automatic processing, criteria for automaticity, and key experiments (Schneider & Shiffrin) that contrast automatic versus controlled processing using visual scanning tasks.
🗂️ Topics Covered
The lecture begins with a recap of filter and capacity models of attention. It then defines automatic processing, its advantages and disadvantages, and criteria for automaticity (Posner & Snyder, 1975). It presents Shiffrin & Schneider’s perspective that automaticity is a matter of degree. The core experimental section describes Schneider & Shiffrin’s visual array scanning studies, contrasting same-category vs. different-category conditions, and discusses implications for attentional load and capacity.
📝 Lecture Summary
RECAP OF LAST LESSONS
Attention can be explained by bottleneck or filter models as well as by capacity models. Filter models assume that interference in attention is caused by a filter or an attenuator. Some filter models assume that filters occur early, before recognition. Others assume that filter occurs late, after semantic analysis. Capacity models assume that interference is caused by overload on the capacity.
AUTOMATICITY
Automatic Processing Tasks vary considerably in the amount of mental effort required to perform them. Some skills become so well-practiced and routine that they require very minimal mental capacity. Cognitive psychologists use the term automatic processing to refer to such skills. The more a process has been practiced, the less attention it requires, and there is speculation that highly practiced processes require no attention at all — such processes are referred to as automatic.
Pros and cons
- It allows us to perform routine activities without much concentration or mental effort; it does not require much attention.
- Automatic processes complete themselves without conscious control by the subject.
- We may make silly mistakes.
- We may fail to remember what we did.
- We are not able to show others how we do a task.
When is a skill automatic? According to Posner and Snyder (1975), a skill is automatic if it:
- Occurs without intention
- Does not give rise to conscious awareness
- Does not interfere with other mental activities
Automaticity: Another Perspective Shiffrin & Schneider (1977) argue that it is best to think of automaticity as a matter of degree rather than a distinct category. A nice demonstration of the way practice affects attentional limitations is the study reported by Underwood (1974) on the psychologist Moray, who has spent many years studying shadowing (split attention studies). Moray can report most of the unattended channels; for him, shadowing has become automatic. Through a great deal of practice, the process of shadowing has become partially automated.
More experiments Schneider & Fisk (1982), Schneider & Shiffrin (1977), Shiffrin & Dumais (1982), and Shiffrin & Schneider (1977) performed a series of experiments contrasting controlled vs. automatic processing.
They said automatic processes complete themselves without conscious control by the subject. In the visual and auditory report tasks reviewed earlier, the registering of the stimuli in sensory memory is an automatic process. Many aspects of driving a car and comprehending language appear to be automatic. Controlled processing seems to require conscious control. Many higher cognitive processes, such as performing mental arithmetic, are controlled.
In their experiments, subjects were required to scan visual arrays. The subjects were given a target letter or number and instructed to scan a series of visual displays for the target.
Visual arrays
1st condition:
- J (letter to be recognized)
- GK
- MF
2nd condition:
- 8 (number to be recognized)
- MN
- L 8
The task Two factors are varied:
- Frame size: Each frame has one, two, or four characters on it.
- Relationship between the target item and the items on the frames: In the same category condition, the target is a letter and all the characters on the frame are letters. In the different category condition, the target is a number surrounded by letters.
If the target appears on the frame, subject responds "yes"; if it doesn't appear on the frame, the subject responds "no".
Results
- In the different category condition: Reaction Time was 80 milliseconds, and Accuracy was 95%.
- In the same category condition: Reaction Time was 400 milliseconds, and Accuracy was also 95%.
In the different category condition: No effect of frame size.
In the same category condition: Accuracy and RT deteriorated as frame size increased.
Schneider and Shiffrin argued that before coming into the laboratory, subjects were so well practiced at detecting a number among letters that this process was automatic. In contrast, when subjects had to identify a letter among letters, controlled processing was needed. In this situation, subjects had to attend separately to each letter in the frame and compare it with the target. All these steps took time, and thus subjects were able to inspect each frame properly and achieve respectable levels of performance only when slides were presented slowly. Also, the more letters that were in a frame, the more slowly the frames had to be presented, since subjects had to check each letter in the frame separately. In contrast, subjects could check all items simultaneously in the different category situation to see if any were numbers. They were able to perform this processing simultaneously because the process was automatic.
💡 Why this matters: This experiment provides direct evidence that automatic processing does not consume attentional capacity, while controlled processing is capacity-limited and slows down with increased load.
Implications Schneider & Shiffrin argued that detecting numbers among letters didn’t place any load on the capacity and that is why subjects were not affected by frame size. Detecting letters among letters, however, was a hard task which placed a lot on attention capacity, which was affected negatively as frame size increased.
⭐ Key Takeaways
Automatic processing requires minimal attention, occurs without intention or conscious awareness, and does not interfere with other mental activities—but can lead to errors and poor recall. Automaticity is best viewed as a matter of degree, not a distinct category, as demonstrated by Moray’s shadowing expertise. In Schneider & Shiffrin’s visual scanning experiments, detecting a number among letters (different category) was automatic, yielding 80ms RT with no frame-size effect, while detecting a letter among letters (same category) required controlled processing, yielding 400ms RT and performance that deteriorated with larger frame sizes. The critical finding is that automatic processes place no load on attentional capacity, whereas controlled processes demand conscious attention and are capacity-limited.
🧠 Quick Revision Questions
- What are the three criteria for a skill to be considered automatic, according to Posner and Snyder (1975)?
- In Schneider & Shiffrin’s experiments, what were the reaction times and accuracy rates for the different category condition vs. the same category condition?
- Why did frame size affect performance in the same category condition but not in the different category condition?
- How does Shiffrin & Schneider’s perspective on automaticity differ from viewing it as a distinct category?
- What is the key difference between automatic processing and controlled processing in terms of attentional capacity?
📘 Lecture 12 — Automaticity (continued)
📖 Overview: This lecture continues the exploration of automatic processing by examining Shiffrin and Schneider's (1977) experiment demonstrating how practice can make discrimination tasks automatic. It introduces Hasher & Zacks' (1979) five criteria for distinguishing between automatic and effortful processes, providing a framework for understanding how different cognitive processes function and are affected by various conditions.
🗂️ Topics Covered
The lecture begins with Shiffrin and Schneider's second experiment showing automatic processing after 2100 trials with specific letter sets. It then examines the implications of automaticity for attention and multitasking. The main body introduces Hasher & Zacks' five criteria for automaticity: intentional vs. incidental learning, effects of instruction and practice, task interference, depression or high arousal, and developmental trends. The lecture concludes with a summary table comparing automatic and effortful processes across all five criteria.
📝 Lecture Summary
Experiment
Shiffrin and Schneider (1977) conducted another experiment where the target always came from one set of letters (B C D F G H J K L) and the distracters were from another set (Q R S T V W X Y Z). After 2100 trials, subjects achieved the same level of accuracy and reaction time as the different condition in the previous experiment. This demonstrated that subjects needed 2100 trials of practice before discriminating between two different sets of letters had become as automatic as discriminating numbers from letters.
📌 Example: In the experiment, after 2100 trials, subjects could identify if a letter was a target (from set B-L) or a distractor (from set Q-Z) automatically.
- Results: Reaction Time = 80 ms, Accuracy = 95%
💡 Why this matters: This shows that automaticity is not just about whether two categories are different (like numbers vs letters), but also requires extensive practice to achieve with similar categories.
Implications
The results demonstrate that processes can become automatic with enough practice. When they do, devoting attention to them is no longer necessary. Performance is no longer affected by the number of processes being performed simultaneously.
Five Criteria for Automaticity
Hasher & Zacks (1979) proposed five criteria to distinguish between automatic and controlled or effortful processes. They also made predictions based on these five criteria.
1. Intentional vs. incidental learning
Intentional learning occurs when we are deliberately trying to learn; incidental learning occurs when we are not deliberately trying to learn. For example, teachers tell children to do things but do not do them themselves, or parents tell children not to lie but lie themselves — children learn from parents' and teachers' actions rather than their words.
🔑 Definition — Incidental learning: Learning that occurs without deliberate intention or awareness of learning.
Incidental learning is as effective as intentional learning for automatic processes but is less effective for effortful learning. For example, we know that in Urdu the letter "seen" occurs more often than the letter "zhe" without trying to learn this information.
2. Effect of instruction & Practice
Instructions on how to perform a task and practice on the task should not affect automatic processes because they can already be carried out very efficiently. For instance, an expert cricketer who comes to the ground and the coach says "when you see the ball, hit it" — this instruction is not effective because the expert already knows what to do.
Both instruction and practice should affect effortful processes. Practice helps in learning well.
3. Task interference
Automatic processes should not interfere with each other because they require little or no capacity.
Effortful processes require considerable capacity and should interfere with each other when they exceed the amount of available capacity.
4. Depression or High arousal
Emotional states such as depression or high emotional arousal can reduce the effectiveness of effortful processes. If we are in a sad mood and someone gives us a difficult and demanding task, we cannot concentrate on that task and cannot learn well. We are not able to learn and pay attention in the classroom when we are in a sad mood.
Automatic processes should not be affected by emotional states. For example, if we have to brush our teeth, we can do it even if we are in a sad mood.
5. Developmental trends
Automatic processes show little change with age. Once a task is practiced, then age does not matter. They (most of the automatic processes) are acquired early and do not decline in old age.
Effortful processes show developmental changes. They are not performed as well by young children or the elderly. There are many things that old people cannot do because these tasks are not practiced. For example, if we are teaching math to an old man, he cannot learn well and easily.
We have to pay attention to all tasks — attention and practice make things automatic. For example, a student who has the habit of reading can succeed even if he is not intelligent.
Summary Table: Differences between Automatic and Effortless Processes
| Criterion | Automatic | Effortful |
|---|---|---|
| Intentional vs. incidental | No difference | Intentional better |
| Effect of instructions and practice | No effects | Improve performance |
| Task interference | No interference | Interference |
| Depression or high arousal | No effects | Poor performance |
| Developmental trends | None | Poor performance |
⭐ Key Takeaways
The lecture establishes that automaticity develops through extensive practice, as demonstrated by Shiffrin and Schneider's experiment requiring 2100 trials for letter discrimination to become automatic. Once processes become automatic, they require no attention, show no interference with other tasks, and are unaffected by emotional states or age. The five criteria from Hasher & Zacks provide a systematic framework for distinguishing automatic from effortful processes: intentional vs incidental learning, effects of instruction and practice, task interference, emotional arousal effects, and developmental trends. The fundamental difference is that automatic processes operate efficiently without capacity limitations, while effortful processes require capacity, show interference, and are affected by instructions, emotional states, and age. For exam purposes, remember that automatic processes show no difference between intentional and incidental learning, no effects from instruction/practice, no task interference, no effects from depression/arousal, and no developmental changes.
🧠 Quick Revision Questions
- How many trials of practice were needed in Shiffrin and Schneider's (1977) experiment for letter discrimination to become automatic, and what were the final reaction time and accuracy results?
- According to Hasher & Zacks' first criterion, how does incidental learning compare to intentional learning for automatic versus effortful processes?
- Why do automatic processes not interfere with each other, while effortful processes do interfere?
- How do emotional states like depression or high arousal differently affect automatic and effortful processes?
- What are the five criteria proposed by Hasher & Zacks for distinguishing between automatic and effortful processes?
📘 Lecture 13 — Automaticity (continued)
📖 Overview: This lecture explores the concept of automaticity in cognitive psychology, focusing on how humans automatically encode frequency information and how selective attention tasks can predict real-world performance. It also examines the paradoxical effects of thought suppression and strategies to manage unwanted thoughts.
🗂️ Topics Covered
This lecture covers the automatic encoding of frequency information, using selective attention tasks to predict flight performance and road accidents, and an in-depth investigation of thought suppression. The lecture details experiments on selective listening as a predictor of pilot and driver safety, and Wegner's white bear experiments on the rebound effect of suppressed thoughts, including remedies for this effect.
📝 Lecture Summary
Judging frequency
Hasher & Zacks (1984) proposed that people are inherently good at judging the relative frequency of events. They argued that this information is encoded automatically without conscious effort. This automatic knowledge allows us to develop expectancies about the world around us.
🔑 Definition — Automatic Encoding: The process by which information, such as the frequency of events, is stored in memory without conscious intent or effort. 💡 Why this matters: This skill is foundational for learning and adapting to our environment, as it helps us predict what is likely to happen.
Predicting flight performance
Gopher & Kahneman (1971) investigated the role of selective attention in pilot training. They noted that flight cadets often failed because they could not appropriately divide their attention or were slow to recognize crucial signals on unattended channels. They tested 100 cadets in the Israeli Air Force using a selective listening task. In this task, two different messages were presented to each ear through headphones. A high tone signaled the right ear as relevant, and a low tone signaled the left ear. Subjects had to report all the digits from the relevant ear.
📌 Example: Three groups of cadets were tested: Group 1 (17 cadets rejected early on light aircraft), Group 2 (41 cadets rejected during jet training), and Group 3 (42 cadets who had reached advanced jet training). The results showed the number who made 3 or more errors:
- Group 1: 76%
- Group 2: 56%
- Group 3: 24% The selective listening task was the best predictor of flight performance, suggesting it is more appropriate than other tests for recruitment. 💡 Why this matters: This has implications for fighter pilot training in the Pakistan Air Force, as a simple listening test can identify candidates with better attentional control.
Predicting Road Accidents
Kahneman, Ben-Ishai & Lotan (1973) studied bus drivers to predict accident rates. They classified drivers into three groups:
- Accident prone drivers: Had 2 or more severe accidents in one year.
- Accident-free drivers: Had no accidents in the same period.
- Intermediate drivers: Fell between the other two groups. The selective listening task had a high correlation with driver safety; those who performed best on the task were safe drivers with a low rate of accidents.
📌 Example: In a different experiment, Mihal & Barett (1976) used seven tests to predict accident involvement of commercial drivers. They found the selective listening task to be the best predictor. This was surprising because a visual task was not as good a predictor. This is likely because the selective listening task is a general measure of attention; those good at switching attention in an auditory task are also good at visual tasks.
Thought suppression
Wegner and colleagues (1987) studied thought suppression to investigate attention to internal sources of information. The classic "white bear" experiment involved two situations:
- Suppression: "Don't think of a white bear." Subjects had to ring a bell whenever they thought of a white bear.
- Expression: "Think of a white bear." Subjects also rang a bell for each thought of a white bear.
📌 Results: The number of bell rings per 5-minute block was recorded.
- Suppression before Expression: Rings started at 3.5 and dropped to 1.
- Suppression after Expression: Rings started at 4.4 and dropped to 1.
- Expression before Suppression: Rings started at 4.5 and dropped to 1.8.
- Expression after Suppression: Rings started at 4.5 and increased to 5.2.
📄 Implications: This demonstrates a paradoxical effect of thought suppression: it produces a preoccupation with the suppressed thought. Subjects use environmental cues to help with suppression, which become associated with the thought. It is better to work on suppression in an environment different from one's usual environment.
Remedies for rebound
The rebound effect (the increased frequency of a thought after suppression) can be reduced.
- Competing thought: Instructions to think about a red Volkswagen instead of a white bear reduced white bear thoughts during expression (Wegner et al., 1987).
- Changing surroundings: Subjects who changed their surroundings had fewer white bear thoughts (Wegner & Schneider, 1989).
⭐ Key Takeaways
The lecture emphasizes that attention is not only crucial for processing external information but also for managing internal thoughts. Judging frequency is an automatic process that helps build expectancies. Selective attention, as measured by a simple listening task, is a powerful predictor of real-world performance in complex tasks like flying and driving. Thought suppression has a paradoxical effect, leading to a rebound of the suppressed thought, but this can be mitigated by using a competing thought or changing one's environment.
🧠 Quick Revision Questions
- According to Hasher & Zacks, how do people encode information about the frequency of events, and why is this useful?
- In Gopher & Kahneman's flight cadet experiment, what task was used as a predictor, and which group of cadets made the fewest errors?
- What was the surprising result of Mihal & Barett's study on road accidents, and what was the proposed explanation?
- Describe the "paradoxical effect" of thought suppression as demonstrated by Wegner's white bear experiment.
- What are two specific remedies for the rebound effect of thought suppression?
📘 Lecture 14 — Pattern Recognition
📖 Overview: This lecture explores how the human perceptual system recognizes visual information, specifically letters and patterns. It contrasts two major theoretical models—template matching and feature analysis—and explains why human recognition is far more flexible than simple machine matching, addressing the core question of perception.
🗂️ Topics Covered
The lecture begins by framing pattern recognition as a critical question for perception theory, then examines Template Matching Models with figures and machine examples, followed by a discussion of Human Flexibility that highlights problems with templates, and finally introduces the Feature Analysis model as a superior alternative with advantages over template matching.
📝 Lecture Summary
Template Matching Models
The template-matching theory assumes that a retinal image of an object is faithfully transmitted to the brain, where it is compared directly to various stored patterns called templates. The perceptual system tries to compare the letter to templates it has for each letter and reports the template that gives the best match. The input represents retinal cells, and the template pattern specifies which retinal cells should be activated for a match.
🔑 Definition — Template: A stored pattern in the brain that an incoming retinal image is compared against to achieve recognition.
Examples in machines include finger-print matching machines that match fingerprints with criminal fingerprints, bar code reading machines in shopping stores, and ATM/credit card reading machines that match card numbers with passwords. However, these machines recognize exact matches—any minor change means there is no match.
💡 Why this matters: The figure illustrates successful matching (a, c), wrong matching (b), wrong part of retina (d), wrong size (e), wrong orientation (f), and non-standard images (g, h). These failures reveal the model's core limitation.
Human flexibility
In Humans we can recognize large letters and small characters, characters in the wrong place and in strange orientations. For example, if you meet your uncle after many years even if he is in a different appearance you can recognize him. This is human flexibility to recognize things, people or places.
Problems with templates
The human perceptual system is much more flexible than template matching. Most patterns we encounter undergo at least some minor changes. Letters are in different font sizes and styles, shapes change slightly, orientations are different—but still we can recognize things.
Feature Analysis model
In this model, stimuli are thought of as combinations of elemental or primitive features. The features for alphabets consist of:
- horizontal lines _
- vertical lines I
- lines at approximately 45 degree angle /
- and curves (
For example, the alphabet T has one vertical line and one horizontal line. All letters are composed of these four patterns. The human mind analyzes letters according to these features.
🔑 Definition — Feature Analysis model: A perceptual theory where stimuli are recognized as combinations of simpler, elemental features rather than whole templates.
How is it better than template model?
Features are mini-templates but features are simpler. One can identify critical features for a letter. For A, two 45-degree angles (/ ) intersect at the top and a horizontal line _ intersects both of these near the middle. The pattern of A consists of these lines plus a specification as to how they should be combined. These features are very much like the output of edge and bar detectors in the visual cortex.
📌 Example: All of the following can be recognized as the same letter A, even in different shapes and different styles: A A A A A A A A A. In the feature model we do not need a template for each letter but only for every feature, which is a great saving. We have 26 letters small and capital. If letters have many features in common, subjects are prone to confuse them. By these four features we can store all 26 letters. This model is also applicable in Urdu letters with the addition of Dot (nukta). Urdu writing is much more complex than English, and Urdu reading is also difficult as well.
The feature model has two main advantages over the template models. First, since the features are simpler, it is easier to see how the system might try to correct for the kinds of difficulties caused by templates model. A second advantage of the feature combination scheme is that it is possible to specify those relationships among features that are most critical to the pattern.
⭐ Key Takeaways
The critical question for perception theory is how sensory information is recognized for what it is. Template matching theory assumes direct comparison to stored templates but fails due to human flexibility in recognizing variations in size, orientation, and style. The feature analysis model overcomes this by using combinations of four primitive features (horizontal lines, vertical lines, diagonal lines, and curves) to recognize all letters, offering simplicity and the ability to specify critical feature relationships. Machine examples like fingerprint and barcode readers highlight the rigidity of template matching, while feature analysis better explains our ability to recognize degraded or varied patterns.
🧠 Quick Revision Questions
- What are the four primitive features used in the feature analysis model for alphabet recognition?
- Name three machine examples of template matching mentioned in the lecture.
- Why does the template matching model fail to explain human pattern recognition?
- How does the feature analysis model handle variations in letter shapes and sizes?
- What is the main advantage of using features over whole templates in pattern recognition?
📘 Lecture 15 — Pattern Recognition Feature Analysis (continued)
📖 Overview: This lecture continues the discussion of feature analysis in pattern recognition, presenting behavioral evidence for features as components in visual recognition. It then extends the feature analysis approach to speech recognition, introducing phonemes and their defining features such as voicing and place of articulation. The lecture demonstrates how feature overlap predicts confusion errors in both visual and auditory perception.
🗂️ Topics Covered
The lecture begins with behavioral evidence for feature analysis from an experiment by Kinney, Marsetta, & Showman (1966), which showed that letters sharing common features are frequently misclassified. It then transitions to speech recognition, highlighting the problem of segmentation and the illusion of word boundaries. The concept of phonemes as basic speech sound units is introduced, followed by a detailed discussion of voicing and place of articulation as critical features. An experiment by Miller & Nicely (1955) demonstrates that phonemes differing by a single feature are more often confused than those differing by multiple features, supporting the feature analysis model.
📝 Lecture Summary
Example
There is fair amount of behavioral evidence for the existence of features as components in pattern recognition. For instance, if letters have many features in common as with C and G, evidence suggests that subjects are particularly prone to confuse them. When such letters are presented for a very brief interval, subjects often misclassify one stimulus as the other.
Kinney, Marsetta, & Showman, (1966) conducted an experiment. In that experiment they presented letters for very brief intervals. The subjects made 29 errors when letter G was presented: 21 involved misclassification as C, 6 misclassification as O, 1 misclassification as B, 1 misclassification as 9. No other errors occurred.
Implications
It is clear that subjects were choosing items with similar features. C, G, O, B, and 9 all share a curve. Such a response pattern would be predicted by a feature analysis model. If subjects can only extract some of the features during a short time, they would have difficulty deciding between letters that share these features.
Speech Recognition or Auditory Recognition
Recognition of speech poses new problems. It is more complex. Segmentation is a major problem because speech is not clearly demarcated the way written text is. Native speakers always mix words together in their speech. They do it unconsciously. It seems there are clear gaps between words but that is an illusion. For example in Urdu: "Kya haal hay" and "kya ho raha hay". The person who does not know Urdu understands these words as "kyaalay" and "kyaoray". The speech appears to be continuous stream of sounds with no obvious word boundaries. It is our familiarity with our own language that leads to the illusion of word boundaries.
Phonemes
Phonemes are the basic vocabulary of speech sounds; it is in terms of them that we recognize. Like, School [s] [k] [u] [l]. S, K, U, L are phonemes. These help us in understanding. Feature analysis and feature combination processes seem to underline speech perception much as they do visual recognition. As with individual letters, individual phonemes can be analyzed as consisting of a number of features. These are given below.
-
Voicing — It is feature or sound of phonemes produced by the vibration of the vocal cords. Like, "Sip" and "zip". [s] is voiceless, [z] is voiced. Place your fingers as you produce each sound.
-
Place of articulation — It refers to the place at which the vocal track is closed or constricted in the production of a phoneme. It is closed at some point in the utterance of most consonants. A bilabial consonant is pronounced using both lips; a consonant pronounced by bringing both lips into contact with each other or by rounding them. In English the bilabials are b, p, m, and w. Labiodentals like f and v are formed because the bottom lip is pressed against the front teeth. Alveolar consonants like t, d, s, z, n, l are formed because the tongue presses against the alveolar ridge of the gums just behind the upper front teeth.
Experiment
Miller & Nicely (1955) presented consonants in noise. Voiced and voiceless pairs included: Bilabial [b] and [p]; Alveolar [d] and [t]. [b] is voiced and [p] is voiceless. [d] and [t] are identical in case of articulation. They presented letters with noise.
Results: Subjects exhibited confusion and reported hearing one sound when actually another sound had been presented. Experimenters were interested in which sound was confused with which. The feature analysis model would predict more confusion with sounds that differed by only a single feature. When presented with [p], subjects more often thought that they heard [t] than that they heard [d]. The phoneme [t] differs from [p] only in terms of place of articulation, whereas the [d] differs in both place of articulation and voicing. Similarly subjects presented with [b] more often thought they heard [p] than [t].
💡 Why this matters: This experiment provides strong empirical support for the feature analysis model in auditory perception, showing that the number of shared features directly predicts the likelihood of confusion errors.
⭐ Key Takeaways
The lecture provides compelling behavioral evidence that feature analysis underlies both visual pattern recognition and speech perception. In vision, letters sharing many features (like curves in C, G, O, B, and 9) are frequently confused under brief presentations. In speech, the fundamental unit of recognition is the phoneme, which can be analyzed in terms of features such as voicing and place of articulation. The critical finding from Miller & Nicely (1955) is that phonemes differing by only a single feature are more often confused than those differing by multiple features, directly supporting the feature analysis model. A key challenge in speech recognition is the segmentation problem, as speech is a continuous stream with no clear word boundaries. Students must remember that our familiarity with our native language creates the illusion of word boundaries, and that both visual and auditory pattern recognition rely on detecting and combining features.
🧠 Quick Revision Questions
- What pattern of errors did Kinney, Marsetta, & Showman (1966) find when they presented the letter G for a brief interval, and how does this support feature analysis?
- Why is segmentation a major problem in speech recognition? Provide an example from Urdu to illustrate the illusion of word boundaries.
- What are the two key features of phonemes discussed in this lecture, and how do they differ between the sounds [p] and [b]?
- In the Miller & Nicely (1955) experiment, why were subjects more likely to confuse [p] with [t] than with [d]?
- How does the feature analysis model explain the confusion errors observed in both visual letter recognition and auditory phoneme recognition?
📘 Lecture 16 — Pattern Recognition (continued)
📖 Overview: This lecture continues the study of pattern recognition, focusing on speech perception. It introduces the concept of voice-onset time as a key acoustic feature distinguishing voiced from unvoiced consonants and describes a landmark adaptation experiment that provided evidence for feature detection in speech perception. The lecture emphasizes how cognitive psychology investigates the underlying processes of perceptual errors.
🗂️ Topics Covered
This lecture defines voice-onset time as the critical interval between air release and vocal cord vibration that differentiates consonants like /b/ and /p/. It then details the adaptation paradigm experiment by Eimas & Corbit (1973), which demonstrated that repeated exposure to a voiced sound shifts perception toward the unvoiced category, supporting feature analysis models. The lecture concludes by linking these findings to broader themes in cognitive psychology, including the importance of studying errors to understand cognitive processes, and discusses template matching and tone in speech recognition.
📝 Lecture Summary
Voice-Onset Time
Voice is human sound and onset means beginning. In pronouncing consonants like [b] and [p], two events occur: closed lips open to release air, and the vocal cords begin vibrating (voicing). For voiced consonants like [b], the release of air and voicing are nearly simultaneous. For unvoiced consonants like [p], the release occurs 60 milliseconds before vibration begins. The perception of a voiced versus unvoiced consonant therefore depends on detecting the presence or absence of a specific interval between release and voicing.
🔑 Definition — Voice-Onset Time: The period of time between the release of air (when closed lips open) and the beginning of vocal cord vibration (voicing). For voiced sounds like /ba/, release and voicing are simultaneous; for unvoiced sounds like /pa/, there is a 60 ms gap.
📐 Formula: Voice-Onset Time = Time(Voicing Onset) – Time(Release Onset) → For /b/: approximately 0 ms; for /p/: approximately 60 ms
📌 Example: In an experiment, participants were presented with the words anbil and anpil to identify the voice timing difference between /b/ and /p/. In the sound wave diagram, the first arrow marks the release, the second arrow marks the voicing. The difference measured was 60 ms for /p/ (long interval) and nearly 0 ms for /b/ (short interval).
Adaptation paradigm
Eimas & Corbit (1973) conducted an experiment called the adaptation paradigm, also known as the fatigue paradigm. The logic is that when we listen to a sound repeatedly, our perceptual system becomes fatigued and expects a different sound to come next.
Eimas and Corbit had subjects listen to repeated presentations of da (voiced). This sound involves a voiced consonant, [d]. They then presented subjects with a series of artificial sounds that spanned the acoustic continuum between ba (voiced) and pa (voiceless). Subjects had to indicate whether each artificial stimulus sounded more like ba or more like pa. Results showed that subjects who, under normal conditions, would report ba as ba were now reporting it as pa.
The researchers reasoned that repeated voicing causes the perceptual system to adapt to voicing and therefore expect unvoiced stimuli. This experiment is critical because it tells us that it is not the sound itself, but rather its features that are being detected.
💡 Why this matters: This finding provides strong evidence for feature analysis models of perception—our perceptual system analyzes stimuli by detecting specific features (like voicing), rather than matching whole sounds to stored templates.
Basic features of feature analysis model
Cognitive psychology is trying to find the causes of perceptual mistakes, because behind these mistakes often lies a theory or model. When we cannot see the problems, we cannot find the weaknesses of systems. Therefore, we cannot explain any system without understanding its process.
In the model of hearing, researchers discuss template matching. There are thousands of sounds in the world. For example, different sounds made by a tonga wala (horse-cart driver) have meanings in other languages. Sounds vary across languages; a small feature can change the entire meaning of a sentence or word. Some words vary with tone.
Cognitive psychology is not just about laboratory experiments—it relates to our daily life. Speech recognition is also concerned with its features.
⭐ Key Takeaways
The most critical concepts from this lecture are voice-onset time as the specific 60 ms interval distinguishing voiced from unvoiced consonants, and the adaptation paradigm experiment which showed that repeated exposure to a voiced sound shifts perception toward the unvoiced category, proving that features (not whole sounds) are detected. Understanding these feature-based perceptual mechanisms helps explain why we make certain speech perception errors and supports feature analysis models over simple template matching. The lecture underscores that cognitive psychology studies real-world processes—speech recognition involves detecting small feature differences that can completely change meaning.
🧠 Quick Revision Questions
- What is voice-onset time, and how does it differ for voiced vs. unvoiced consonants?
- In the Eimas & Corbit (1973) experiment, what was the stimulus presented repeatedly, and what shift in perception did subjects show afterward?
- What does the adaptation paradigm reveal about whether the brain detects whole sounds or specific features?
- How does template matching differ from feature analysis in models of hearing?
- According to the lecture, why is studying perceptual errors important for cognitive psychology?
📘 Lecture 17 — PATTERN RECOGNITION (continued)
📖 Overview: This lecture continues the topic of pattern recognition by introducing the Gestalt theory of perception. It explains how the human mind organizes sensory information into unified wholes, demonstrating that the whole is perceived differently from the sum of its parts, and how slight differences in organization can lead to vastly different interpretations.
🗂️ Topics Covered
The lecture covers the Gestalt theory of perception, introduced by German psychologists in the 20th century, and explains how principles of perceptual organization determine how we segment objects. It uses multiple diagrams to show how scattered shapes are arranged into unified patterns, how we perceive rows and columns differently, and how ambiguous figures like the young/old lady are interpreted. The lecture emphasizes that a slight difference in organization changes our visual perception and that the mind tends to complete incomplete figures, proving that the unified whole is different from the sum of its parts.
📝 Lecture Summary
Gestalt Theory of Perception
Pattern recognition is constructing a new thing. Various principles determine how we segment an object into components. Only after segmentation does perceptual pattern matching come into play. Ideas for aggregating various lines and images into segments are very similar to what have been referred to as gestalt principles of perceptual organization. Gestalt is a new concept presented by German Psychologists in the 20th century.
In the following diagram, on the right side there are different figures and shapes. And on the left side these shapes and figures are arranged in bicycle shape. All infrastructures are scattered in left side but on the other side these all are arranged in unified whole (pattern) and giving meanings. We see things as whole.
🔑 Definition — Gestalt principles: Principles of perceptual organization that determine how we segment objects and see them as unified wholes, where the whole is different from the sum of its parts.
The text presents several diagrams of circles that can be perceived as rows or columns. The number of circles remains the same but in a slightly different shape or organization, and people see them differently. In all examples, the key word is organization.
📌 Example: An ambiguous figure is shown where some see a picture of a young lady, while others see a picture of an old lady. This happens because of different perceptual organization.
The central idea of gestalt principles is the unified whole is different from the sum of its parts.
Perceptual Organization and Boundaries
A diagram shows lines that some perceive as a pattern broken into two parts, while others perceive two kinds of lines: short and long. A boundary line is drawn between lines.
In another diagram, the dimension is changed through shading and shifting of position, showing these are different patterns. This line and pattern convert the concept of one pattern into a concept that these are two different patterns. The purpose of all these diagrams is to show how a little difference in the same things can be viewed differently. In all diagrams, parts are the same but their organization is different, and this slight difference in organization changes our entire visual perception. A slight difference causes huge perceptual difference.
💡 Why this matters: Understanding how minimal changes in organization dramatically alter perception is crucial for designing visual displays, user interfaces, and understanding visual illusions in cognitive psychology.
Completing Incomplete Figures
Two diagrams show the same lines with a slight difference of line, so our perceptual system perceives them differently.
We perceive four different patterns of circles and also see a square in the middle of circles. Even though there is no square, we perceive it.
In one diagram where we perceived a square, there is exactly a square, but we are also showing four circles behind the square. In reality, these are not circles — they are curves. This is a significant proof of Gestalt theory. The mind usually doesn't like to see things incomplete. We tend to complete figures that are incomplete, and the mind sees it. We don't do it consciously.
Look at a figure showing black circles with grey shading in a few circles. In the following figure, the circles are arranged in such a way that the grey shading makes an imaginary rectangle. In the next figure, the rectangle has shaded dark. We perceive this figure as 8 circles and one rectangle, but in reality, there is only one complete circle at the bottom. This all means we see something else — our visual information receives incomplete information, but our mind tends to complete it.
📌 Example: A figure showing an incomplete triangle and three dots is perceived as a six-pointed star, because we are used to seeing stars. In the figure of a real star, we can see two rectangles as well.
🔑 Definition — Perceptual completion: The tendency of the mind to complete incomplete figures and perceive them as whole patterns, a key principle of Gestalt theory.
⭐ Key Takeaways
The central idea of Gestalt theory is that the unified whole is different from the sum of its parts, meaning we perceive patterns as organized wholes rather than individual components. A slight difference in organization causes a huge perceptual difference, demonstrating how our visual system is highly sensitive to arrangement. The mind actively completes incomplete figures, proving that perception is not just passive reception of sensory data but an active constructive process. Gestalt principles explain ambiguous figures (like the young/old lady) where the same physical stimulus can be perceived differently based on perceptual organization. Understanding these principles is essential for grasping how pattern recognition works and how the brain constructs meaningful patterns from sensory input.
🧠 Quick Revision Questions
- What is the central idea of Gestalt theory regarding the relationship between the whole and its parts?
- Why do different people perceive the same ambiguous figure (young/old lady) differently?
- In the diagram where circles arranged with grey shading appear to form a rectangle, what principle explains why we perceive a rectangle that isn't actually complete?
- What does the example of perceiving a six-pointed star from an incomplete triangle and three dots demonstrate about perception?
- How does a slight difference in the organization of the same elements lead to huge perceptual differences?
📘 Lecture 18 — PATTERN RECOGNITION (continued)
Gestalt Theory of Perception
📖 Overview: This lecture continues the discussion of pattern recognition by delving into Gestalt theory of perception. It explains how the human mind naturally organizes visual elements into meaningful wholes, covering key principles like proximity, similarity, and closure. The lecture also connects these principles to object perception through influential models by Palmer, Marr & Nishihara, and Biederman.
🗂️ Topics Covered
The lecture begins with an introduction to Gestalt theory, using visual examples of figures like a cross or square to demonstrate how perception completes missing information. It then covers the six main Gestalt principles of organization: proximity, similarity, good continuation, closure, good form, and symmetry. The concept of figure-ground perception is introduced using the example of Queen Elizabeth’s vase. The lecture also discusses Palmer’s (1977) experiment on recognition of parts, and concludes with object perception models by Marr & Nishihara (1978) and Biederman (1987).
📝 Lecture Summary
Gestalt Theory of Perception
Gestalt principles are basic principles of perception that explain how we organize visual input. For example, a figure made of four lines (two horizontal and two vertical) is often perceived as a cross, even though it could be seen otherwise. It is a human tendency to complete a figure even when some information is missing, such as perceiving an incomplete set of lines as a square.
🔑 Definition — Gestalt Theory: A theory of perception that emphasizes the organization of visual elements into whole, meaningful patterns.
Gestalt Principles of Organization
The most common Gestalt principles of organization are:
- Proximity: Items that are close together in space or time tend to be perceived as belonging together or forming an organized group.
- Similarity: Similar items tend to be organized together; same things are considered one thing.
- Good continuation: The tendency to perceive a line that starts in one way as continuing in the same way.
- Closure: Perceptual processes that organize the perceived world by filling in gaps in stimulation.
- Good Form: A type of closure where we fill in the gaps to perceive a complete form rather than disconnected lines.
- Symmetry: The tendency to organize things to make a balanced or symmetrical figure that includes all parts.
📌 Example: The text provides examples of these principles but no specific numeric example.
Queen Elizabeth’s Vase
The concept of figure and ground is crucial in Gestalt theory. Queen Elizabeth’s vase is a gift given to her on her silver jubilee. It is a vase, but we perceive it as a figure of two faces. We constantly change background into figure and figure into background.
🔑 Definition — Figure-Ground: A key concept in visual perception where we distinguish an object (figure) from its background (ground).
Palmer (1977)
Palmer (1977) studied subjects' recognition of figures. He first showed subjects a stimulus (a) and then asked them to decide whether fragments (b)-(e) were part of the original figure. Stimulus (a) tends to organize itself into a triangle and a bent letter 'n'. Palmer found that subjects could recognize the parts most rapidly when they were the segments predicted by the Gestalt principles. Recognition depends critically on the initial segmentation of the figure.
💡 Why this matters: This experiment demonstrates that how we initially break down a visual scene (segmentation) directly impacts how quickly and accurately we can recognize its components.
Object Perception
The basic idea of object perception is that a familiar object can be seen as a known configuration of simple components.
Marr & Nishihara (1978) proposed that familiar objects can be seen as configurations of simple pipe-like components. For example, a human model can be made from different types and sizes of cylinders. We perceive these cylinders as a human model. Similarly, a diagram of a bird is also made of cylinders. This is a computer model.
Biederman (1987) proposed that there are three stages in recognition of an object as a configuration of simpler components:
- Segmentation into sub-objects
- Classify the category of each sub-object
- Recognition as a pattern made of sub-objects
⭐ Key Takeaways
The most critical points from this lecture are the six Gestalt principles of organization (proximity, similarity, good continuation, closure, good form, symmetry) which explain the human tendency to perceive organized wholes. The concept of figure-ground shows how perception can shift between an object and its background. Palmer's experiment proved that recognition speed depends on initial segmentation based on Gestalt principles. Finally, object perception models by Marr & Nishihara (using cylinders) and Biederman (three-stage recognition) explain how we recognize complex objects as configurations of simpler components.
🧠 Quick Revision Questions
- List the six Gestalt principles of organization explained in this lecture.
- What is the difference between closure and good form?
- In Palmer's (1977) experiment, which fragments were recognized most rapidly by subjects?
- What is the fundamental concept demonstrated by Queen Elizabeth's vase?
- According to Biederman (1987), what are the three stages of object recognition?
📘 Lecture 19 — Object Perception (continued)
📖 Overview: This lecture continues the discussion of object perception, focusing on Biederman’s (1987) theory of recognition-by-components. It explains how objects are recognized by segmenting them into simpler sub-objects called geons, classifying those sub-objects, and then recognizing the object as a pattern made from those components. The lecture also covers experimental evidence supporting this theory.
🗂️ Topics Covered
The lecture presents Biederman’s three-stage model of object recognition: segmentation into sub-objects, classification of sub-object categories using geons, and recognition of the object as a pattern. It then describes the experimental evidence comparing segment deletion versus component deletion, showing that brief presentations favor recognition of figures with component deletion while longer presentations favor segment deletion.
📝 Lecture Summary
Object Perception (continued)
Biederman (1987) proposed three stages in recognition of an object as a configuration of simpler components:
- Segmentation into sub-objects
- Classify the category of each sub-object
- Recognition as a pattern made of sub-objects
1. Segmentation
Sub-objects are defined by their line contours – Edge and bar detectors from David Marr. Hoffman and Richards (1985) – Gestalt principles can be used to segment an outline representation of an object into sub-objects. They observe that where one segment joins another there is typically a concavity in the line outline.
2. Classification of sub-object categories
Second, once an object has been segmented into basic sub-objects, one can classify the category of each sub-object. Biederman (1987) argues that there are 36 basic categories of sub-objects, which he calls Geons – abbreviation of geometric ions.
The five geons are shown in the lecture figures. There are geons that make different objects. For example, geon number 1 is a part of a telephone, and geon number 5 is a receiver of a telephone set. We can vary the size of the shape and get different objects. The same type of geons makes different things through different organization. Altogether, Biederman proposes there are 36 geons that can be generated in this manner and that they serve as an alphabet for composing objects, much as letters or phonemes serve as the alphabet for building up words.
🔑 Definition — Geons: 36 basic categories of sub-objects (geometric ions) that serve as the fundamental building blocks for object recognition.
3. Recognition of object
Third, having identified the pieces out of which the object is composed and their configuration, one recognizes the object as the pattern composed from these sub-objects or pieces. Thus, recognizing an object is like recognizing a letter; the sub-objects become the features.
Minor details and variations don’t matter. As in the case of letter recognition, there are many small variations on the underlying features (geons) that should not be critical for recognition. Edges are more important than texture to define geons. Color, texture, and small detail should not matter. This predicts that schematic line drawings of complex objects which allow the basic geons to be identified should be recognized as quickly as detailed color photographs of the objects.
💡 Why this matters: This prediction means that our visual system prioritizes structural edges over surface details, which has practical implications for designing effective visual displays, diagrams, and even how we process information in daily life.
Biederman’s Experiment
Biederman conducted an experiment to test his hypothesis. He showed different shapes to subjects and created two conditions:
- Segment deletion
- Component deletion
In this experiment, some objects had whole components deleted while others had all the components present but segments of these components were deleted. They presented these two types of degraded figures to subjects for various brief intervals and asked them to identify the objects.
For example, in the lecture figures, the components of elephants were separated and shown. In one figure the components were deleted.
Biederman’s Evidence
The critical assumption is that object recognition is mediated by recognition of the components of the object. The results showed, at very brief presentations (65-100 milliseconds), subjects were more accurate at the recognition of figures with component deletion than segment deletion. This reversed for the longer 200 milliseconds presentations.
Biederman reasoned that at the very brief intervals, subjects were not able to identify the components with segment deletion and so had difficulty in recognizing the objects. With 200 milliseconds exposure, however, sub-objects were able to recognize all the components in either condition. Since there were more components in the condition with segment deletion, they had more information as to object identity.
So, we can conclude that we do not split reality into geons or anything else. We bring all the information into our sensory store. There is a difference between reality and representativeness, reality and perception, and reality and recognition.
🔑 Definition — Segment deletion: Removing parts of the edges/contours within each component while keeping all components present. 🔑 Definition — Component deletion: Removing entire sub-objects/geons from the figure.
📐 Key Finding: At 65-100ms → Component deletion better recognized. At 200ms → Segment deletion better recognized. 📌 Example: An elephant figure with its trunk entirely removed (component deletion) was recognized more easily at very brief exposures than a figure with the trunk present but with gaps in its outline (segment deletion). At longer exposures, the figure with more complete components (segment deletion) provided more total information for identification.
⭐ Key Takeaways
The most critical points from this lecture are: Biederman’s recognition-by-components theory proposes that objects are recognized through three stages: segmentation, classification of geons, and pattern recognition. There are 36 basic geons that serve as an alphabet for composing objects, much like letters form words. Edges are more important than texture, color, or small details for defining geons, which predicts that line drawings should be as recognizable as color photographs. The experimental evidence shows that at very brief exposures (65-100ms), component deletion leads to better recognition than segment deletion, but this reverses at longer exposures (200ms). This demonstrates that component identification is critical for initial object recognition, while overall information becomes more useful with longer viewing time.
🧠 Quick Revision Questions
- What are the three stages of object recognition according to Biederman (1987)?
- How many basic geons does Biederman propose exist, and what function do they serve?
- Why are edges more important than texture, color, or small details for object recognition?
- In Biederman’s experiment, why were figures with component deletion recognized more accurately at very brief presentations (65-100ms) compared to segment deletion?
- What was the critical reversal in results between short (65-100ms) and longer (200ms) presentation times?
📘 Lecture 20 — Attention & Pattern Recognition
📖 Overview: This lecture explores the critical role of attention in combining features to perceive patterns, and how context and prior knowledge influence pattern recognition. It explains the difference between single feature detection and feature conjunction, introduces top-down processing, and discusses the Word Superiority Effect, demonstrating that word context can enhance letter perception.
🗂️ Topics Covered
The lecture covers the role of attention in feature integration through Treisman & Gelade's experiment on detecting a T among distractors, demonstrating that feature conjunction takes longer than single feature detection. It then introduces context and pattern recognition using ambiguous figures, explains top-down processing as high-level knowledge guiding perception, and concludes with the Word Superiority Effect and its explanation by Rumelhart & Siple regarding inferential perception.
📝 Lecture Summary
ATTENTION & PATTERN RECOGNITION
Attention is required to combine features to perceive patterns. Treisman & Gelade (1980) conducted an experiment where subjects tried to detect a T in an array of 30 I's and Y's. Subjects could detect this by looking for the single cross-bar feature of the T, taking about 800 milliseconds. When subjects had to detect a T in an array of 30 I's and Z's, they could not use a single feature; they had to look for the conjunction of features (both the vertical and horizontal bars), taking more than 1200 milliseconds. This difference of about 400 milliseconds shows that feature conjunction requires more attention than single feature detection. The researchers varied display size: subjects showed little difference between conditions for displays with fewer than five letters, but with more distracters, attention became substantially overloaded. This demonstrates that for familiar letters, the deficit in perceiving feature conjunctions only appears with large displays. 💡 Why this matters: This explains why we can automatically recognize single features but need focused attention to combine features in complex visual scenes.
Context & Pattern Recognition
The lecture presents an example of ambiguous figures where the same picture is perceived as an animal when placed in a line of animal pictures (hen, rabbit, dog, cat) and as a human when placed in a line of human pictures (man, woman, child, girl). This demonstrates that context influences how we perceive patterns—the identical image is interpreted differently based on surrounding stimuli.
Top-Down Processing
When context or general world knowledge guides perception, this is called top-down processing, because high-level general knowledge determines the interpretation of low-level perceptual units. An example is the phrase "THE CAT" where the middle letter in the word can be seen as either A or H depending on the surrounding letters. This shows that the surrounding words have an effect on our perception of individual letters.
🔑 Definition — Top-Down Processing: When high-level general knowledge and context guide the interpretation of low-level perceptual units.
📌 Example: In "THE CAT", the middle letter is ambiguous—it can be read as 'A' in "CAT" or 'H' in "THE"—showing that the word context determines which letter we perceive.
Word Superiority Effect
The Word Superiority Effect (WSE) was demonstrated by Reicher and Wheeler (1970). They presented a brief display of either a letter (e.g., D) or a word (e.g., WORD). Immediately afterward, subjects were given a pair of alternatives and instructed to report which they had seen. If shown D, alternatives might be D or K; if shown WORD, alternatives might be WORD or WORK. Subjects were about 10% better in the word condition—they more accurately discriminated between D and K in the context of a word than as letters alone.
Why WSE?
Rumelhart & Siple (1974) explained why subjects are more accurate in the word condition. Suppose subjects identify the first three letters as WOR. There are several four-letter words consistent with WOR beginning: WORK, WORD, WORM, WORN, WORT. If subjects detect only the bottom curve in the fourth letter, in the word context they need only detect one feature to perceive the fourth letter (e.g., the curve of D distinguishes it from K). However, when the letter is presented alone and subjects detect the curve, they will not know whether the letter was B, D, O, or Q, since each is consistent with the curve feature. When the letter is presented alone, they must identify a number of features. This analysis implies that perception is inferential—like the curve in D helps in recognition only when combined with word context.
🔑 Definition — Word Superiority Effect (WSE): The phenomenon where letters are more accurately recognized when presented in the context of a word than when presented alone.
📐 Formula/Mechanism: Word context → fewer features needed to identify a letter (e.g., one feature in WORD) vs. letter alone → multiple features needed (e.g., curve in D could be B, D, O, or Q).
📌 Example: Subjects correctly identified D vs. K about 10% more often when D appeared in "WORD" (as opposed to "WORK") than when D was presented alone, because the "WOR" context limited possibilities.
⭐ Key Takeaways
The critical takeaway is that attention is essential for combining single features into complex patterns, as shown by Treisman & Gelade's experiment where feature conjunction detection took 400ms longer and was more affected by display size than single feature detection. Context powerfully shapes perception, as demonstrated by ambiguous figures and the top-down processing principle, where high-level knowledge determines interpretation of low-level units. The Word Superiority Effect shows that word context enhances letter recognition by about 10%, because perception is inferential—a word context reduces the number of features needed for identification compared to isolated presentation. Students must understand that these effects reveal perception as an active, constructive process rather than a passive registration of sensory input.
🧠 Quick Revision Questions
- What was the difference in reaction time between detecting a T by a single feature (among I's and Y's) versus detecting a T by feature conjunction (among I's and Z's) in Treisman & Gelade's experiment?
- Why did display size affect subjects' performance more in the conjunction condition than in the single feature condition?
- What is top-down processing, and how does the "THE CAT" example illustrate it?
- What is the Word Superiority Effect, and how much better were subjects in the word condition compared to the letter-alone condition?
- According to Rumelhart & Siple, why does word context make letter recognition easier than recognizing a letter alone?
📘 Lecture 21 — Pattern Recognition (Continued)
📖 Overview: This lecture examines neural networks as a model for pattern recognition, focusing on McClelland and Rumelhart's interactive activation model. It explains how parallel distributed processing (PDP) resolves the paradox of bottom-up versus top-down processing, demonstrating how features, letters, and words interact through excitatory and inhibitory connections to facilitate word recognition.
🗂️ Topics Covered
The lecture covers the interactive activation model by McClelland and Rumelhart, which solves the bottom-up versus top-down paradox in pattern recognition. It explains neural network components including nodes, links, excitatory and inhibitory connections, weights, activation rules, and learning rules. The lecture also discusses how PDP models have improved computer functioning and how word superiority effect emerges from network interactions.
📝 Lecture Summary
Pattern Recognition (Continued)
This model is also called PDP's or Parallel Distributed Processing. McClelland and Rumelhart (1981) made a pattern recognition network that solves the paradox of bottom-up versus top-down processing. They implemented this network to model our use of word structure to facilitate recognition of individual letters. In this model, individual features are combined to form letters and individual letters are combined to form words. This is a connectionist model that depends heavily on excitatory and inhibitory activation processes. Activation spreads from the features to excite the letters, and from the letters to excite the words. Alternative letters and words inhibit each other. Activation can also spread down from the words to excite the component letters, so a word can support the activation of a letter and promote its recognition. In such a system, activation will tend to accumulate at one word and repress the activations of other words through inhibition. The dominant word will support the activation of its component letters, and these letters will repress the activation of alternative letters. The word superiority effect is due to the support a word gives to its component letters. The computation proposed by McClelland and Rumelhart's interactive activation model is extremely complex, as is the computation of any model that stimulates neural processing. This process helps us in understanding how neural processing underlies pattern recognition.
The network diagram shows a simpler version of a neural network displaying the word WORK. At one level there are feature detectors (not shown). At another level there are letter detectors. We recognize W because of features; O, R, and K are recognized similarly. All words activate because of letters. Four words of WORK are activated — these three words are competitors in word recognition. WORK is recognized by four letters. This is a parallel distributor processing model, also called a neural network model of pattern recognition. It is called a neural network because there is an abstract concept or quantity called nodes.
🔑 Definition — Connectionist model: A model of pattern recognition where individual features combine to form letters, and letters combine to form words, using excitatory and inhibitory activation processes. 🔑 Definition — Word superiority effect: The phenomenon where a word supports the activation of its component letters, making letter recognition easier within a word context than in isolation. 📌 Example: In the word WORK, the word-level activation supports recognition of each individual letter (W, O, R, K) while inhibiting alternative letters at the same positions.
Neural Networks
Neural networks consist of: Nodes, Links (both excitatory and inhibitory), Weights, and Learning which consists of re-adjustment of weights.
Nodes are a set of processing units. Nodes should not be confused with neurons. Nodes are represented by features, letters, and words in the interactive activation model. They can acquire different levels of activation. All boxes in the diagram are nodes; lines are links. Nodes are connected through these lines.
Patterns of connections: Nodes are connected to each other by excitatory or inhibitory connections that differ in strength. Another important concept is activation rules, which specify how a node combines its excitatory and inhibitory inputs with its current state of activation.
🔑 Definition — Nodes: A set of processing units in a neural network that can acquire different levels of activation; they are represented by features, letters, and words in the interactive activation model.
Excitatory connections are those connections that make other nodes active. Nodes connected with excitatory connections are active or charged. Inhibitory connections are those connections that make other nodes relax and switch off. Because of these connections the neural network exists.
State of Activation: Nodes can be activated to various degrees. We become conscious of nodes that are activated above a threshold level of conscious awareness. We become aware of letter K in the word WORK when it receives enough excitatory influences from feature and word levels.
A Learning Rule: Learning generally occurs by changing the weights of the excitatory and inhibitory connections between the nodes. The Learning rule specifies how to make these changes in the weights. Two aspects are initial weights and re-adjustment of weights.
🔑 Definition — Excitatory connections: Connections between nodes that make other nodes active or charged. 🔑 Definition — Inhibitory connections: Connections between nodes that make other nodes relax and switch off. 🔑 Definition — Activation rules: Rules that specify how a node combines its excitatory and inhibitory inputs with its current state of activation. 🔑 Definition — Learning rule: A rule that specifies how to change the weights of excitatory and inhibitory connections between nodes to enable network improvement. 💡 Why this matters: The learning component is the most important feature of a neural network model because it enables the network to improve its performance. In a lab in California, a computer learned how to speak by reading and re-reading simple English sentences — improving from its own mistakes.
PDP and its significance
Parallel processing models have improved computer functioning, leading to super computers, also called parallel computers. Multiple processors that communicate with each other work faster than serial processing computers. The paradoxes are resolved — like forests being seen at the same time as the trees, words are seen at the same time as the letters. Context helps in object perception; object perception helps in perception of context.
⭐ Key Takeaways
The interactive activation model by McClelland and Rumelhart resolves the bottom-up versus top-down paradox by allowing activation to spread bidirectionally between features, letters, and words. Neural network components include nodes (processing units), excitatory connections (which activate other nodes), inhibitory connections (which suppress other nodes), and weights that determine connection strength. The learning rule enables networks to improve performance by re-adjusting weights based on experience. The word superiority effect demonstrates how word-level information supports letter recognition through top-down activation. PDP models have led to parallel computers that process information faster than serial processing systems.
🧠 Quick Revision Questions
- What are the three levels of processing units in McClelland and Rumelhart's interactive activation model for word recognition?
- How do excitatory and inhibitory connections differ in their function within a neural network?
- What is the word superiority effect and how does the interactive activation model explain it?
- What is the role of the learning rule in neural network models, and what does it specify?
- How does parallel distributed processing resolve the paradox of bottom-up versus top-down processing?
📘 Lecture 22 — Pattern Recognition (Continued)
📖 Overview: This lecture explores how sentence context and speech context influence word recognition and pattern perception. Through classic experiments by Tulving, Mandler & Baumal (1964) and Warren (1970), it demonstrates that prior context reduces the amount of bottom-up information needed to identify words, and even allows listeners to "fill in" missing phonemes, showing that perception is actively shaped by top-down expectations.
🗂️ Topics Covered
The lecture covers the effect of sentence context on word identification through the Tulving, Mandler & Baumal (1964) experiment, the interaction between bottom-up information and context, the implications of context for filling in missing words, the phoneme restoration effect demonstrated by Warren (1970), and the Warren & Warren (1970) follow-up study showing how subsequent words determine identification of ambiguous speech sounds.
📝 Lecture Summary
Effects of Sentence Context
Cognitive psychologists study whether sentence context affects word recognition. An experiment by Tulving, Mandler & Baumal (1964) demonstrated this effect at the multiword level. They used sentences like "Countries in the United Nations form a military alliance" and "The huge slum was filled with dirt and disorder." Each sentence provided an eight-word context preceding a critical word. In various conditions, subjects saw 0, 4, or 8 words of context before the critical word (e.g., "disorder") was presented for a very brief duration ranging from 0 to 140 milliseconds. The experimenters manipulated both the amount of context and the exposure duration to study how bottom-up information interacted with context.
🔑 Definition — Bottom-up information: Information derived directly from sensory input (e.g., visual features of letters or acoustic features of speech). 🔑 Definition — Top-down processing: The use of prior knowledge, expectations, or context to influence perception.
📐 Formula: Probability of correct identification = f(context amount, exposure duration) → The chance of correctly identifying a word increases both when more context is provided and when the word is shown for longer.
📌 Example: Results showed: At 0ms flash duration, 0 context = 0% correct, 4 context = 10%, 8 context = 16%. At 140ms flash duration, 0 context = 70%, 4 context = 80%, 8 context = 98%. At 60ms flash duration, 0 context = 30%, 4 context = 60%, 8 context = 70%. The maximum effect of context was seen at 60ms exposure. The effect diminished somewhat between 60 and 140ms because subjects in the 8-word context condition performed almost perfectly and showed little benefit from further exposure, whereas subjects in the 0-word condition continued to benefit from longer exposure. These results indicate that subjects can take advantage of context to improve their identification of words.
Implications
This experiment shows that we can use sentence context to help identify words. With context, we need to extract less information from the word itself in order to identify it. We can also use context to fill in words that didn't even occur. We are able to fill in missing letters as we read the sentence, perhaps not even noticing they were missing. PDP (Parallel Distributed Processing) models can help explain this effect of context better than other models while also accounting for feature analysis.
💡 Why this matters: Context reduces the sensory evidence required for recognition, meaning perception is not purely data-driven but also expectation-driven.
Context and Speech
Text works this way, but does it work with speech? Speech is experienced sequentially in a more linear fashion than text. The Phoneme Restoration Effect was demonstrated in an experiment by Warren (1970). He had subjects listen to the sentence: "The state governors met with their respective legislatures convening in the capital city." A 120 ms pure tone replaced the middle 's' in "legislatures." Only 1 in 20 subjects reported hearing the pure tone, and even that person wasn't able to locate it clearly.
🔑 Definition — Phoneme Restoration Effect: A phenomenon in speech perception where listeners "fill in" a missing or replaced phoneme based on context, often not noticing the manipulation.
📌 Example: Warren & Warren (1970) presented subjects with sentences such as:
- "It was found that the *eel was on the axle" → subjects reported hearing "wheel"
- "It was found that the *eel was on the shoe" → subjects reported hearing "heel"
- "It was found that the *eel was on the orange" → subjects reported hearing "peel"
- "It was found that the *eel was on the table" → subjects reported hearing "meal" In each case, the asterisk denotes a phoneme replaced by non-speech. The identification of the critical word was determined by subsequent context.
Implications
The implications of the phoneme restoration experiments are: (1) Context fills in gaps and affects our perception just as in texts. (2) The identification of the critical word is determined by what comes after the critical word (e.g., "heel," "peel," "meal," "wheel" are critical words). (3) Thus, the identification of words can depend on the perception of subsequent words. In a nutshell, when you face a problem, you should grasp the context. When you grasp the context, you are able to understand and handle the problem.
⭐ Key Takeaways
Context powerfully influences both visual word recognition and speech perception by reducing the amount of bottom-up sensory information needed for identification. The Tulving, Mandler & Baumal experiment shows that increasing sentence context and exposure duration both improve word identification, with the largest contextual benefit occurring at intermediate exposure durations. The phoneme restoration effect demonstrates that listeners use context to perceptually fill in missing or distorted speech sounds, often without noticing the manipulation. Crucially, word identification can depend on information from words that come after the target word, showing perception is actively constructed from both prior and subsequent context. PDP models provide a useful framework for explaining how top-down context and bottom-up feature analysis interact in pattern recognition.
🧠 Quick Revision Questions
- What was the main independent variable manipulated in the Tulving, Mandler & Baumal (1964) experiment besides exposure duration?
- At which exposure duration was the maximum effect of sentence context observed on word identification?
- What is the phoneme restoration effect, and what does it reveal about speech perception?
- In the Warren & Warren (1970) experiment, what determined whether subjects heard "wheel," "heel," "peel," or "meal"?
- How do PDP models help explain the effect of context on pattern recognition compared to other models?