eng507 — Midterm Summary (Lectures 1–22)
📘 Lecture 1 — Introduction to the Course-I
📖 Overview: This lecture provides an introduction to the course Phonetics and Phonology (ENG507), outlining its aims, objectives, and evaluation criteria. It establishes the foundational understanding of why the study of speech sounds is crucial within the broader field of linguistics and introduces the basic division of human speech sounds into consonants and vowels.
🗂️ Topics Covered
The lecture introduces the overall structure of the course, beginning with a general introduction to Phonetics and Phonology as a branch of linguistics. It then explains the importance of studying speech sounds for language documentation and typological studies, identifies English (specifically Received Pronunciation) as the focus language, and details the course aims and evaluation criteria. Finally, it provides a basic introduction to the two main categories of speech sounds: consonants and vowels.
📝 Lecture Summary
Topic-001: Introduction to the Course
Phonetics and Phonology (P&P), with course code ENG507, is an introductory level module focused on the linguistic study of speech sounds. The course has two main components: the phonetic component, which provides a detailed analysis of speech sounds with an emphasis on articulatory phonetics (how sounds are physically produced), and the phonology component, which examines the internal structure of words, both simple (simplex) and complex (complex word forms). The module is designed for beginners and requires students to complete regular problem-solving exercises from the main textbook by Ladefoged and Johnson.
Topic-002: Introduction to the Course (Why studying phonetics and phonology?)
Linguistics is the study of language, involving various building blocks. Phonetics and phonology is the branch of linguistics that deals with human speech sounds, describing sounds like vowels (monophthongs and diphthongs) and consonants. This field studies sets of phonemes and sound patterns, which can be dynamic (as in connected speech) or static (as in isolation). Expertise in this area is crucial for describing undocumented spoken languages, aiding in language documentation and language description, as well as for typological studies and cross-linguistic comparisons.
Topic-003: Focus Language - English
The primary focus language for this course is English, specifically based on the Received Pronunciation (RP) , which is the British accent. The course will describe English sounds in detail, covering the features of consonants, vowels, and diphthongs. Additionally, examples from local languages such as Urdu, Punjabi, Sindhi, and Pashto will be included to provide a greater understanding of relevant topics and facilitate comparisons.
Topic-004: Aims and Objective of the Course
Upon completing the course, students will be able to understand how human sound is produced, know the physical properties of human sounds, and study the suprasegmental features and features of connected speech. They will also gain a greater awareness of IPA symbols and be able to transcribe any kind of English text. Furthermore, the course prepares students for advanced coursework in Experimental Phonology and teaches them about modern software used in phonological research.
Topic-005: Evaluation Criteria for the Course
The standard evaluation criteria of the university will be used, including quizzes, assignments, and a Graded Discussion Board (GDB) . The course will have a mid-term exam and a final-term exam covering all important topics. Transcription will also be a significant component of the evaluation.
Topic-006: Introduction to Vowels and Consonants
Human speech sounds are divided into two broad categories: consonants and vowels. A consonant is a speech sound where the air is at least partly blocked, while a vowel is a sound where there is no obstruction and the air passes freely. Consonants are classified by their places of articulation, manners of articulation, and voicing. Vowels are classified by the position of the tongue, the part of the tongue used, and lip-rounding. Vowels are further classified into pure vowels (monophthongs) and diphthongs.
🔑 Definition — Phonetics and Phonology: The branch of linguistics that deals with the linguistic study of human speech sounds from a phonetic and phonological perspective. 🔑 Definition — Articulatory Phonetics: The study of how speech sounds are physically produced by the vocal organs. 🔑 Definition — Received Pronunciation (RP) : The standard, prestigious British accent of English used as the model for describing English sounds in this course. 🔑 Definition — Consonant: A speech sound in which the airflow is at least partly blocked. 🔑 Definition — Vowel: A speech sound in which there is no obstruction and air passes freely through the vocal cavity.
⭐ Key Takeaways
The most critical points from this lecture are that Phonetics and Phonology (ENG507) is an introductory course in the linguistic study of speech sounds, focusing on articulatory phonetics (how sounds are made) and phonological structure (how sounds pattern in words). The course uses English (Received Pronunciation) as its focus but also draws on local languages for comparison, and it aims to teach students about sound production, IPA transcription, and suprasegmental features. Evaluation is based on standard university criteria including quizzes, assignments, GDB, and exams, with transcription as a key component. Finally, the foundation of the subject is the distinction between consonants, where airflow is blocked, and vowels, where it flows freely.
🧠 Quick Revision Questions
- What are the two main components of the Phonetics and Phonology (ENG507) course?
- Define the difference between a consonant and a vowel in terms of airflow.
- What is the focus language for this course, and which specific accent is used as a model?
- Name two fields of study for which expertise in phonetics and phonology is important, as mentioned in the lecture.
- What are the three criteria used to classify a consonant?
📘 Lecture 2 — Introduction to the Course-II
📖 Overview: This lecture provides a foundational overview of the English sound system, introducing the classification of vowels and consonants in the RP (BBC) accent. It then distinguishes between the related but distinct fields of phonetics and phonology, explaining what each discipline studies and why it matters for understanding language.
🗂️ Topics Covered
The lecture introduces English vowels (20 total: 12 pure vowels and 8 diphthongs), English consonants (24 total: plosives, nasals, fricatives, affricates, and approximants), and provides the IPA transcription for all 44 sounds of RP English. It then introduces phonology as the study of sound systems and contrastive sounds in a particular language, before introducing phonetics as the study of human speech sounds, including its three major branches.
📝 Lecture Summary
Topic-007: Introduction to English Vowels
The English RP (BBC) accent has 44 sounds in total. Out of these, 20 are vowels, which are divided into two main categories: pure vowels (monophthongs) and diphthongs. There are 12 pure vowels, further split into 5 long vowels and 7 short vowels.
🔑 Definition — Pure vowels (monophthongs): Vowel sounds produced with a single, stable articulatory position, without gliding from one sound to another.
🔑 Definition — Long vowels: Vowels held for a longer duration in pronunciation (marked with the length symbol ː).
🔑 Definition — Short vowels: Vowels produced with a shorter duration than long vowels.
📌 Example (short vowels): ɪ as in pit, e as in pet, æ as in pat, ʌ as in putt, ɒ as in pot, ʊ as in put, ǝ as in another.
📌 Example (long vowels): iː as in bean, ɑː as in barn, ɔː as in born, uː as in boon, ɜː as in burn.
Topic-008: Introduction to English Diphthongs
Diphthongs are vowel sounds that involve a glide from one vowel quality to another within the same syllable. English has 8 diphthongs, which are categorized into two types: centering diphthongs (which end with the ǝ sound, gliding towards the center of the mouth) and closing diphthongs (which end with either ɪ or ʊ sounds, gliding towards a closer, higher position).
🔑 Definition — Diphthongs: Vowel sounds that consist of a movement or glide from one vowel articulation to another.
🔑 Definition — Centering diphthongs: Diphthongs that glide towards the central vowel ǝ.
🔑 Definition — Closing diphthongs: Diphthongs that glide towards either the close front vowel ɪ or the close back vowel ʊ.
📌 Example (centering diphthongs): ɪǝ as in peer, eǝ as in pair, ʊǝ as in poor.
📌 Example (closing diphthongs ending in ɪ): eɪ as in bay, aɪ as in buy, ɔɪ as in boy.
📌 Example (closing diphthongs ending in ʊ): ǝʊ as in no, aʊ as in now.
Topic-009: Introduction to English Consonants
English consonants are classified by their manner of articulation. There are five main categories, totaling 24 consonant sounds. These categories include plosives (6 sounds), nasals (3 sounds), fricatives (9 sounds), affricates (2 sounds), and approximants (4 sounds).
🔑 Definition — Plosives: Consonant sounds produced by completely blocking the airflow and then releasing it with a burst.
🔑 Definition — Nasals: Consonant sounds produced by lowering the velum to allow air to escape through the nose.
🔑 Definition — Fricatives: Consonant sounds produced by forcing air through a narrow channel, creating friction.
🔑 Definition — Affricates: Consonant sounds that begin as a plosive and release as a fricative.
🔑 Definition — Approximants: Consonant sounds where the articulators approach each other but do not narrow the vocal tract enough to cause turbulent airflow.
📌 Example (plosives): p as in pin, b as in bin, t as in tin, d as in din, k as in kin, g as in gum.
📌 Example (nasals): m as in sum, n as in sun, ŋ as in sung.
📌 Example (fricatives): f as in fine, v as in vine, θ as in think, ð as in this, s as in seal, z as in zeal, ʃ as in sheep, ʒ as in measure, h as in how.
📌 Example (affricates): ʧ as in chain, ʤ as in Jane.
📌 Example (approximants): l as in light, r as in right, w as in wet, j as in yet.
Topic-010: IPA Transcription of English Sounds
The IPA (International Phonetic Alphabet) provides a standardized set of symbols for representing the sounds of spoken language. For RP English, the 44 sounds are transcribed using the specific IPA symbols listed below.
📌 Example (Full IPA inventory of RP English):
- Long vowels:
iː,ɑː,ɔː,uː,ɜː - Short vowels:
ɪ,e,æ,ʌ,ɒ,ʊ,ǝ - Diphthongs:
eɪ,aɪ,ɔɪ,ǝʊ,aʊ,ɪǝ,eǝ,ʊǝ - Consonants (plosives):
p,b,t,d,k,g - Consonants (nasals):
m,n,ŋ - Consonants (fricatives):
f,v,θ,ð,s,z,ʃ,ʒ,h - Consonants (affricates):
ʧ,ʤ - Consonants (approximants):
l,r,w,j
Topic-011: Introduction to Phonology
Phonology is the study of the sound systems of a particular language. A key concern in phonology is whether sounds are contrastive—that is, whether substituting one sound for another changes the meaning of a word. For example, in English, [r] and [l] are contrastive because "road" and "load" have different meanings. Phonologists describe the full set of contrastive consonants and vowels (the phonemes) in a sound system. Their study also extends to larger units like syllables, phrases, rhythm, tone, and intonation of a specific language.
🔑 Definition — Phonology: The study of how sounds function in a particular language or languages, focusing on contrastive patterns and systems.
💡 Why this matters: Understanding phonology explains why certain sound differences (like [r] vs [l]) matter for meaning in English, while others might not, and it forms the basis for analyzing sound patterns in any language.
Topic-012: Introduction to Phonetics
Phonetics, as a discipline, is the general study of human speech sounds. It is concerned with the physical properties of sounds, including how they are produced, transmitted, and perceived. This includes understanding how sounds are articulated using the mouth, nose, teeth, and tongue, and how the ears hear those sounds and distinguish them. Phoneticians may also analyze the physical properties of sounds, such as their waveform, using computer programs like Praat. There are three major branches of phonetics: articulatory phonetics (how sounds are produced), acoustic phonetics (the physical properties of sound waves), and auditory phonetics (how sounds are perceived by the ear and brain).
🔑 Definition — Phonetics: The scientific study of human speech sounds, their production, transmission, and perception. 🔑 Definition — Articulatory phonetics: The branch of phonetics concerned with how speech sounds are produced by the vocal organs. 🔑 Definition — Acoustic phonetics: The branch of phonetics concerned with the physical properties of speech sounds as sound waves. 🔑 Definition — Auditory phonetics: The branch of phonetics concerned with how speech sounds are perceived by the listener's ear and brain.
⭐ Key Takeaways
There are 44 sounds in English RP, consisting of 20 vowels (12 pure vowels: 5 long and 7 short; plus 8 diphthongs: 3 centering and 5 closing) and 24 consonants (grouped into plosives, nasals, fricatives, affricates, and approximants). Vowel length is contrastive, marked by the length symbol ː for long vowels. Diphthongs are categorized by their endpoint—centering diphthongs end in ǝ, while closing diphthongs end in ɪ or ʊ. Phonology studies contrastive sounds within a specific language's system, whereas phonetics is the general scientific study of all human speech sounds across three main branches: articulatory, acoustic, and auditory. The IPA provides a standard transcription system for all 44 English sounds.
🧠 Quick Revision Questions
- How many total sounds are there in English RP, and what is the breakdown between vowels and consonants?
- What is the difference between a pure vowel (monophthong) and a diphthong?
- Name the three categories of English consonants according to manner of articulation, and give one example sound for each.
- Explain the key difference between the fields of phonetics and phonology.
- Why are the sounds [r] and [l] important examples in English phonology?
📘 Lecture 3 — Introduction to Key Concepts in Phonetics and Phonology (P&P)-I
📖 Overview: This lecture establishes the foundational conceptual framework for the entire course by defining and differentiating the core subfields of phonetics and phonology. It introduces the critical distinctions between phones, phonemes, and allophones, and provides a detailed survey of the three main branches of phonetics, explaining their unique foci and methodologies. Understanding these concepts is essential for analyzing how speech sounds are produced, organized, and perceived.
🗂️ Topics Covered
The lecture begins by contrasting phonetics (the study of sound production) with phonology (the study of sound organization in language). It then introduces and distinguishes three key terms: phone, phoneme, and allophone. Following this, the lecture details the three major branches of phonetics: articulatory, acoustic, and auditory phonetics, explaining what each branch studies and why it is important.
📝 Lecture Summary
Topic-013: Phonetics vs. Phonology
Phonetics and phonology are both important subfields of linguistics that deal with speech sounds and overlap each other. The key difference is that phonology is the study of how sounds are organized in individual languages. It focuses on the organization of sounds by studying speech patterns (e.g., phonological rules within a specific language). The key words for describing phonology are ‘distribution’ and ‘patterning’ related to speech. Phonologists may look into questions like – why there is a difference in the plurals of cat and dog; the former ends with an ‘s’ sound, whereas the latter ends with the ‘z’ sound. Phonetics, on the other hand, is the study of the actual process of sound making. Phonetics has been derived from the Greek word ‘phone’ meaning sound or voice. It covers the domain of speech production and its transmission and reception.
Topic-014: Introduction to Key Concepts in Phonetics and Phonology
There are various terms which are frequently used in phonetics and phonology, mainly including phone, phoneme, and allophone. A phone is a sound (or a segment) which has some physical feature and the term is mostly used in a non-technical sense. A phoneme is the smallest meaningful unit of sound (therefore, the smallest unit in phonology) in a language, and this meaningful unit of sound is one that will change one word into another word. For example, the difference in both ‘white’ and ‘right’ (ignore spellings, focus on sounds) is the difference of sounds (/w/ – /r/) which are phonemes, and they have the ability to change meaning. Similarly, take another example of ‘cat’ vs. ‘bat’ (/k/ – /b/). Linguists have also defined a phoneme as a group or class of sound events having common patterns of articulation. An allophone is a definable systematic variant of a phoneme. Compare the following sets: the ‘s’ sound in words like sill, still, and spill; the ‘k’ sound in words like key and car; the ‘t’ sound in words like true and tea; and the ‘n’ sound in words like tenth and ten. If you carefully analyze these words, you should find that the specific sound is not exactly the same in the given word examples. But since these variants do not change meaning (and we simply take them as alternate sounds), they are called allophones.
🔑 Definition — Phone: A sound (or segment) with a physical feature, used in a non-technical sense. 🔑 Definition — Phoneme: The smallest meaningful unit of sound in a language that can change the meaning of a word (e.g., the /p/ and /b/ in pat and bat). 🔑 Definition — Allophone: A systematic phonetic variant of a phoneme that does not change meaning (e.g., the aspirated [pʰ] in pin vs. the unaspirated [p] in spin). 📌 Example: In English, the /t/ sound in top (aspirated [tʰ]) and the /t/ sound in stop (unaspirated [t]) are allophones of the same phoneme /t/. They are different phones but do not create a meaning difference like the phonemes /t/ and /d/ do in tip and dip.
Topic-015: Types of Phonetic Studies
Phonetics is the scientific study of speech sounds. It has three major branches: articulatory phonetics, acoustic phonetics, and auditory phonetics. The central concerns in phonetics are the discovery of how speech sounds are produced; how they are used in spoken language; how we can record speech sounds with written symbols; and how we hear and recognize different sounds. The second area is where phonetics overlaps with phonology: usually in phonetics we are only interested in sounds that are used in meaningful speech, a field sometimes known as linguistic phonetics. Thirdly, there has always been a need for agreed conventions for using phonetic symbols that represent speech sounds; the International Phonetic Association has played a very important role in this regard. Finally, the auditory aspect of speech is very important: the ear is capable of making fine discriminations between different sounds, so much so that sometimes it is not possible to define in articulatory terms precisely what the difference is, but we can still hear it. Phonetics is a multidisciplinary field, studying language in terms of ‘general linguistics’, ‘language development’, ‘dialectology’, ‘sociolinguistics’, ‘psycholinguistics’, ‘anatomy’, ‘physiology’, ‘developmental psychology’, ‘robotics’, and ‘information processing’.
Topic-016: Articulatory Phonetics
Articulatory phonetics deals with studying the making of single sounds by the vocal tract. It is the branch of phonetics which studies the way in which speech sounds are made (‘articulated’) by the vocal organs. It derives much of its descriptive terminology from the fields of anatomy and physiology and is sometimes referred to as physiological phonetics. The classification of sounds used in the International Phonetic Alphabet (IPA), for example, is based on articulatory variables. Important discussions included in this field are: air stream mechanism, speech production, places of articulation, manners of articulation, phonation (voicing), and other processes such as the oro-nasal process and the description of vowel production. 💡 Why this matters: This branch provides the descriptive foundation for the IPA and is the most accessible entry point for analyzing speech production.
Topic-017: Acoustic Phonetics
Acoustic phonetics is related to the study of physical attributes of sounds produced by the vocal tract. It is the branch of phonetics which studies the physical properties of speech sound as transmitted between mouth and ear according to the principles of acoustics (the branch of physics devoted to the study of sound). It is primarily dependent on the use of instrumental techniques of investigation (such as Praat software), particularly electronics, and some grounding in physics and mathematics is a prerequisite for advanced study of this subject. Acoustic analysis can provide a clear, objective datum for the investigation of speech – the physical ‘facts’ of speech sounds (such as duration, formants F1, F2, and F3, etc.). Thus, acoustic evidence is often referred to when one wants to support an analysis being made in articulatory or auditory phonetic terms. 💡 Why this matters: This branch provides objective, measurable data to support and verify analyses made in other branches of phonetics.
Topic-018: Auditory Phonetics
Auditory phonetics deals with understanding how the human ear perceives sound and how the brain recognizes different speech units. This branch of phonetics studies the perceptual response to speech sounds as mediated by ear, auditory nerve, and brain. It is a very less well-studied area of phonetics, mainly because of the difficulties encountered as soon as one attempts to identify and measure psychological and neurological responses to speech sounds. On the other hand, anatomical and physiological studies of the ear are well advanced, as are techniques for the measurement of hearing, and the clinical use of such studies is now established under the headings of audiology and audiometry. The subject is closely related to studies of auditory perception within the domain of psycholinguistics. 💡 Why this matters: This branch focuses on the "reception" end of the communication chain, which is critical for understanding speech recognition and language disorders.
⭐ Key Takeaways
The single most critical distinction to master is the difference between phonetics (the physical production and properties of sounds) and phonology (the abstract, mental system of sound organization). You must be able to define and give examples for a phone (any physical sound segment), a phoneme (a meaning-distinguishing sound unit, e.g., /p/ vs. /b/), and an allophone (a non-meaning-distinguishing variant of a phoneme, e.g., the aspirated [pʰ] and unaspirated [p]). The three branches of phonetics—articulatory (production), acoustic (physical transmission), and auditory (perception)—each study a different stage of the speech chain. Finally, remember that the phoneme is a theoretical group or class, and its members are the allophones.
🧠 Quick Revision Questions
- What is the primary difference in focus between phonetics and phonology?
- Using the example of the /t/ sound in the words top, stop, and little, identify which are phones, which is the phoneme, and which are allophones.
- What are the three major branches of phonetics, and what does each one study (production, transmission, or perception)?
- According to the lecture, what is the official definition of a phoneme?
- Why is auditory phonetics considered a "less well-studied" area compared to the other two branches?
📘 Lecture 4 — Introduction to Key Concepts in Phonetics and Phonology (P&P)-II
📖 Overview: This lecture continues the foundational exploration of phonetics and phonology by introducing experimental and generative approaches to the field. It then provides a detailed anatomical and physical account of speech production, including the roles of articulators, the mechanics of sound waves, and the critical oro-nasal process that distinguishes oral from nasal sounds.
🗂️ Topics Covered
The lecture covers six main topics: experimental phonetics and phonology, which uses hypothesis-based, quantitative methods across articulatory, acoustic, and auditory fields; generative phonology, a rule-based theory pioneered by Chomsky and Halle; articulatory phonetics, which studies the principal articulators involved in speech; the four components of speech production (airstream, phonation, oro-nasal, and articulatory processes); the physical nature of sound waves and their variations in air pressure; and the oro-nasal process, controlled by the velum, which determines whether a sound is oral or nasal.
📝 Lecture Summary
Topic-019: Experimental Phonetics and Phonology
Experimental phonetics and phonology integrates research in experimental phonetics, experimental psychology, and phonological theory to provide a hypothesis-based investigation of phonological phenomena. While much phonetic work is descriptive (accounting for pronunciation) or prescriptive (stating how sounds ought to be pronounced), an increasing amount is experimental—aimed at developing and scientifically testing hypotheses. Experimental phonetics is quantitative, based on numerical measurement, and relies on controlled experiments to ensure results are caused only by the factor being investigated.
🔑 Definition — Experimental Phonetics: The branch of phonetics that uses controlled, quantitative experiments to test hypotheses about speech production, acoustics, and perception.
Experimental research is now carried out in all fields of phonetics: in the articulatory field, we measure how speech is produced; in the acoustic field, we examine the relationship between articulation and the resulting acoustic signal, looking at physical properties of speech sounds; in the auditory field, we perform perceptual tests to discover how the listener’s ear and brain interpret speech signals. Topics explored include infant speech perception, categorical perception of child acquisition, and changes in perception. Peter Ladefoged’s 1967 work explored stress in respiratory activity, the nature of vowel quality, and perception and production of speech.
Topic-020: Generative Phonology
A major change in phonological theory came in the 1960s when Morris Halle and Noam Chomsky showed that many sound processes observable in phonology are actually regulated by grammar and morphology. This area of phonology relates to specific phonological rules within languages—rules that describe substitutions, deletions, and insertions of sounds in specific contexts. To highlight these rules, an elaborate, algebra-like writing method was evolved, best seen in The Sound Pattern of English (Chomsky and Halle, 1968).
🔑 Definition — Generative Phonology: An approach to phonology based on an abstract, underlying phonological representation of speech that requires rules to convert it into phonetic realizations.
Though this type of phonology became extremely complex and has been largely replaced by newer approaches, many of those approaches are still classed as generative because they are based on the principle of an abstract, underlying phonological representation requiring rules. Theories stemming from generative phonology include autosegmental phonology, metrical phonology, lexical phonology, and optimality theory.
Topic-21: Articulatory Phonetics - I
Articulatory phonetics is the branch of phonetics that studies articulators and their actions related to human speech production. We produce speech sounds by moving parts of our articulators (body parts) through muscle contractions. Most relevant movements take place in the mouth and throat area (though chest activity for breath control is also important). Parts of the mouth and throat that we move when speaking are called articulators. This branch studies the principal articulators—such as the tongue, lips, lower jaw, teeth, velum (soft palate), uvula, and larynx—and other processes related to speech production, including features of vowels and consonants and their specific properties (places and manners of articulation, phonation, etc.).
🔑 Definition — Articulators: The parts of the mouth and throat area that move when speaking, including the tongue, lips, lower jaw, teeth, velum, uvula, and larynx.
Topic-22: Speech Production
The process of speech production mainly includes respiration, phonation, articulation, and resonance. To produce speech, we need the airstream mechanism (to activate speech), the exploitation of the airstream at the larynx (called phonation or voicing), the modification of the air passage with articulators at the cavity (oral or nasal), and finally the transfer of energy. Speech production is a term used for the activity of the respiratory, phonatory, and articulatory systems during speech, along with associated processes for their coordination and use. A contrast is drawn with receptive aspects like speech perception and recognition.
As the anatomy of speech, experts (like Ladefoged) highlight four main components:
- The airstream process – all ways of pushing air out that provide energy for speech
- The phonation process – the actions of the vocal folds
- The oro-nasal process – the possibility of the airstream going out through the mouth (as in [v] or [z]) or the nose (as in [m] and [n])
- The articulatory process – the movements of the tongue and lips interacting with the roof of the mouth and the pharynx
💡 Why this matters: Understanding these four processes is essential for analyzing how any speech sound is physically produced, from the air supply to the final shaping of the sound.
Topic-23: Sound Waves
A sound wave is the pattern of disturbance caused by the movement of energy traveling through air. Sound consists of small variations in air pressure that occur very rapidly one after another. These variations are caused by actions of the speaker’s vocal organs superimposed on the outgoing flow of lung air. For voiced sounds, the vibrating vocal folds chop up the stream of lung air so that pulses of relatively high pressure alternate with moments of lower pressure. These variations move through the air like ripples on a pond; when they reach the listener’s ear, they cause the eardrum to vibrate. A graph of a sound wave is similar to a graph of the eardrum’s movements. Understanding physical features of sound waves—such as amplitude, loudness, and time duration of vibration—is important for many phonetic studies, and sound waves play a key role in acoustics.
🔑 Definition — Sound Wave: The pattern of disturbance caused by energy traveling through air, consisting of rapid variations in air pressure.
Topic-024: The Oro-Nasal Process
The oro-nasal process determines the possibility of the airstream going out through the mouth (as in [v] or [z]) or the nose (as in [m] and [n]). Consider the consonants at the end of rang, ran, ram (ŋ, m, n)—all nasal sounds. When saying these consonants alone, air comes out through the nose. In a sequence, the point of articulatory closure moves forward: from velar in ‘rang’ [ŋ], through alveolar in ‘ran’ [n], to bilabial in ‘ram’ [m]. In each case, air is prevented from going out through the mouth but can go out through the nose because the soft palate (velum) is lowered. In most speech, the soft palate is raised so there is a velic closure. When it is lowered and there is an obstruction in the mouth, we have a nasal consonant. Raising or lowering the velum controls the oro-nasal process, the distinguishing factor between oral and nasal sounds.
🔑 Definition — Oro-Nasal Process: The process controlled by the velum that determines whether the airstream exits through the mouth (oral sounds) or the nose (nasal sounds).
📐 Rule: Velum raised → velic closure → oral sound (air exits through mouth) / Velum lowered + mouth obstruction → nasal sound (air exits through nose) 📌 Example: For the word "rang" [ræŋ], the velum is lowered, air is blocked at the velar place in the mouth but escapes through the nose, producing a nasal consonant.
⭐ Key Takeaways
The lecture establishes that phonology has evolved from descriptive to experimental and generative approaches, with experimental phonetics using controlled, quantitative methods to test hypotheses about speech, and generative phonology introducing abstract underlying representations that require rules to produce phonetic output. Speech production is anatomically broken down into four processes—airstream, phonation, oro-nasal, and articulatory—each contributing a distinct function to the creation of sound. Sound waves are physical patterns of air pressure variation, essential for acoustic analysis, while the oro-nasal process, controlled by the velum, is the precise mechanism that distinguishes between oral and nasal consonants. A student must understand these foundational concepts, including the key articulators, the four speech production components, and how the velum's position determines whether a sound is oral or nasal.
🧠 Quick Revision Questions
- What three fields of experimental phonetics does Ladefoged identify, and what does each field investigate?
- What is the core principle of generative phonology, and how does it differ from earlier phoneme-based approaches?
- List the four main components of speech production as described by Ladefoged, and explain what each component does.
- What is a sound wave, and how does the vibration of the vocal folds contribute to its formation in voiced speech?
- Describe the oro-nasal process: what determines whether a sound is oral or nasal, and what happens to the velum in each case?
📘 Lecture 5 — Articulatory Phonetics-II
📖 Overview: This lecture explains how speech sounds are produced by describing the specific articulatory gestures (places of articulation) and the various manners of articulation that classify consonant sounds. Understanding these concepts is essential for accurately describing and transcribing speech sounds in any language.
🗂️ Topics Covered
This lecture covers the major articulatory gestures (bilabial, labiodental, dental, alveolar, retroflex, palato-alveolar, palatal, and velar) and then systematically explains the manners of articulation. It defines stops (both oral and nasal), fricatives, and approximants, including a discussion of additional consonantal gestures like affricates. Finally, it distinguishes between trills, taps, and flaps as specific consonantal gestures.
📝 Lecture Summary
Topic-025: Articulatory Gestures
To fully describe a speech sound, we need to understand the movements of articulators, called articulatory gestures, where one articulator moves toward another. The major places of articulation for English sounds include:
- Bilabial: Made with two lips (e.g., /p/, /b/).
- Labiodental: Lower lip raises to touch the upper front teeth (e.g., /f/, /v/).
- Dental: Tongue tip or blade contacts the upper front teeth (e.g., the first sound in 'thigh').
- Alveolar: Tongue tip or blade contacts the alveolar ridge (e.g., /t/, /d/, /n/, /s/, /z/, /l/).
- Retroflex: Tongue tip curls back against the back of the alveolar ridge. This is common in Pakistani languages like Urdu, Sindhi, Pashto, Balochi, and Punjabi.
- Palato-alveolar: Tongue blade contacts the back of the alveolar ridge (e.g., /ʃ/ in 'shy').
- Palatal: Front of the tongue contacts the hard palate (e.g., /j/ in 'yes').
- Velar: Back of the tongue contacts the soft palate (e.g., /k/, /g/).
Topic-026: Manner of Articulation
The manner of articulation describes the type and degree of obstruction a sound makes to the airflow. Consonants are classified by their manner into two major types: obstruents (stops, fricatives, affricates) which create a significant obstruction, and sonorants (nasals, liquids, glides) which allow freer airflow. The International Phonetic Association (IPA) classifies consonants by both their manner and place of articulation.
Topic-027: Stop: Oral and Nasal
A stop is a sound produced by a complete closure in the vocal tract, which is then suddenly released. This involves two processes: the closure (the 'stop') and the burst (the release). The outward rush of air upon release is called plosion. Oral stops (plosives) in English are [p, b, t, d, k, g]. A distinction is made between oral stops and nasal stops (e.g., [m, n, ŋ]), where the air is released through the nose.
Topic-028: Fricative
A fricative is made by forcing air through a narrow gap, creating a hissing noise. Fricatives can be voiced (e.g., [z]) or voiceless (e.g., [s]). A distinction is made between sibilant fricatives (strong and audible, like [s, ʃ]) and strident fricatives (weaker, like [θ, f]). BBC pronunciation has nine fricative phonemes: voiceless /f, θ, s, ʃ, h/ and voiced /v, ð, z, ʒ/.
🔑 Definition — Fricative: A consonant produced by forcing air through a narrow constriction, generating a hissing noise.
Topic-029: Approximants
An approximant is a consonant that makes very little obstruction to the airflow. Traditionally, they are divided into two groups:
- Semivowels: Very similar to close vowels but produced as a rapid glide (e.g., [w] in 'wet', [j] in 'yet').
- Liquids: Sounds with an identifiable but non-obstructive constriction, including laterals like [l] in 'lead' and non-fricative [r] as in 'read'. BBC English has four approximant sounds: [l], [r], [w], [j].
Topic-030: Additional Consonantal Gestures
An affricate is a type of consonant that begins as a plosive and is immediately followed by a fricative at the same place of articulation (e.g., [tʃ] and [dʒ] in 'church' and 'judge'). It is often considered a single phoneme in English.
🔑 Definition — Affricate: A sound that begins with a complete closure and ends with a fricative release at the same place of articulation.
Topic-031: Trill, Tap and Flap
These are types of central approximants, distinguished by the number and type of articulatory contact:
- Tap: A single, rapid, up-and-down movement of the tongue tip against the alveolar ridge (e.g., the middle sound in 'pity' in an American accent, transcribed as [ɾ]).
- Flap: A single, rapid, front-and-back movement of the tongue tip, often involving a curling-back motion. It is common in Indo-Aryan languages (e.g., the retroflex [ɽ]).
- Trill: The articulator (e.g., the tongue tip) is set in continuous motion by the air current, striking against the target multiple times (e.g., the Scottish /r/).
⭐ Key Takeaways
You must be able to precisely describe a consonant sound by its place and manner of articulation. The seven primary places of articulation for English are bilabial, labiodental, dental, alveolar, palato-alveolar, palatal, and velar. For manner, the key categories are stops (oral and nasal), fricatives (voiced/voiceless, sibilant/strident), approximants (liquids and glides), and affricates. Finally, you should distinguish between the different types of 'r' sounds: taps (single contact), flaps (curling contact), and trills (continuous contact).
🧠 Quick Revision Questions
- What is the difference between a bilabial and a labiodental sound?
- Name the two major types of stops and give an example sound for each.
- What distinguishes a sibilant fricative from a strident fricative?
- What is an affricate consonant, and what are its two components?
- Distinguish between a tap [ɾ] and a flap [ɽ] in terms of articulator movement.
📘 Lecture 6 — Articulatory Phonetics-III
📖 Overview: This lecture explores the acoustics of consonants through waveforms, delves into the articulation and classification of vowel sounds, distinguishes long vowels from diphthongs, and introduces suprasegmental features. Understanding these concepts is crucial for analyzing speech sounds beyond individual segments and for phonetic transcription.
🗂️ Topics Covered
This lecture covers the waveforms of consonants, detailing how manners of articulation appear in acoustic patterns. It then examines the articulation of vowel sounds, focusing on lip shape and tongue position. The sounds of vowels are defined phonetically and phonologically, including formant frequencies. Long vowels and diphthongs are explained with examples from BBC English. Finally, an introduction to suprasegmental features like pitch, loudness, and stress is provided.
📝 Lecture Summary
Topic-032: The Waveforms of Consonants
The waveforms of consonants display distinctive acoustic characteristics. While places of articulation are not visible, differences in manners of articulation—stop, nasal, fricative, and approximant—are usually apparent. The difference between voiced and voiceless sounds is also visible. For vowels, the lips open and amplitude gets larger. For a stop sound, the closure and burst are easy to judge. A voicing bar appears for voiced sounds, showing small voicing vibrations instead of a flat line. A fricative has a more nearly random waveform pattern.
🔑 Definition — Voicing Bar: Small, low-energy vibrations visible in a waveform that indicate a voiced sound, as opposed to a flat line for voiceless sounds. 📐 Formula: [Not applicable; this is a qualitative observation.] 📌 Example: The waveform of the phrase 'my two boys know how to fish' (Ladefoged and Johnson, 2012, p. 18) shows a voicing bar for the /b/ in 'boys' (voiced) but not for the /t/ in 'two' (voiceless). 💡 Why this matters: Waveform analysis allows phoneticians to visually distinguish different consonant types and voicing, even without hearing the sound.
Topic-033: The Articulation of Vowel Sounds
In vowel articulation, the articulators do not come close together, and the air stream is relatively undisturbed. Vowels make the least obstruction to airflow. They are almost always at the center of a syllable. Each vowel is distinguished by: (1) lip shape: rounded (e.g., /u:/), neutral (e.g., /ə/), or spread (e.g., /i:/); (2) tongue part raised: front (e.g., /æ/ in 'cat'), middle, or back (e.g., /ɑ:/ in 'cart'); and (3) tongue height: close to the roof or low in the mouth. Lip rounding is important in some languages.
🔑 Definition — Articulators: Speech organs (e.g., tongue, lips, jaw) that move to produce different sounds. 📐 Formula: [Not applicable.] 📌 Example: The vowel /i:/ in 'see' involves spread lips, a front tongue raised high and close to the roof of the mouth; the vowel /u:/ in 'too' involves rounded lips, a back tongue raised high.
Topic-034: The Sounds of Vowels
Vowels are defined both phonetically and phonologically. Phonetically, they are articulated without complete closure or audible friction; air escapes evenly over the center of the tongue. If air escapes solely through the mouth, vowels are oral; if some air is released through the nose, they are nasal. Phonetic classification involves (a) lip position (rounded, spread, neutral) and (b) tongue part raised and height. Acoustically, vowels are distinguished by the first two formant frequencies, F1 and F2. F1 is inversely related to vowel height (smaller F1 = higher vowels). F2 is related to front/back position (smaller F2 = more back vowels).
🔑 Definition — Formant Frequencies (F1, F2): Acoustic resonances of the vocal tract that distinguish vowels; F1 correlates with tongue height, and F2 correlates with tongue frontness/backness. 📐 Formula: F1 ∝ 1/vowel height; F2 ∝ vowel frontness (smaller F2 = back vowel). 📌 Example: The vowel /i/ (high, front) has a low F1 and high F2; the vowel /ɑ/ (low, back) has a high F1 and low F2.
Topic-035: Long Vowels and Diphthongs
Long vowels are transcribed with the diacritic [ː] (e.g., /iː/). A contrast of length (short vs. long) is important; stressed syllables are often longer than unstressed. English has phonemic long/short contrasts, e.g., short vowels: /i, e, ɒ, ʊ, ə/ and long vowels: /iː, ɑː, ɔː, uː, ɜː/. A diphthong contains a glide from one vowel quality to another within the same syllable. BBC English has three ending in /ɪ/ (/eɪ, aɪ, ɔɪ/), two ending in /ʊ/ (/əʊ, aʊ/), and three ending in /ə/ (/ɪə, eə, ʊə/).
🔑 Definition — Diphthong: A vowel sound that begins at one vowel quality and glides to another within the same syllable. 📐 Formula: [Not applicable; it is a description of sound change.] 📌 Example: The word 'boy' contains the diphthong /ɔɪ/, starting with a /ɔ/ quality and gliding to /ɪ/. The word 'house' contains the diphthong /aʊ/, starting with /a/ and gliding to /ʊ/.
Topic-036: Introduction to Suprasegmental
The term suprasegmental ('supra' = above, 'segments' = sounds) refers to aspects of sound like intonation that are not properties of individual vowels and consonants. It is used predominantly by American writers; British work prefers the term prosodic. Commonly mentioned suprasegmental features include pitch, loudness, tempo, juncture, syllable, rhythm, and stress.
🔑 Definition — Suprasegmental/Prosodic Features: Aspects of speech (e.g., pitch, stress, rhythm) that extend over multiple segments (vowels and consonants) and affect entire syllables, words, or phrases. 📐 Formula: [Not applicable.] 📌 Example: In the question "You're coming?" vs. the statement "You're coming.", the change in intonation (pitch pattern) is a suprasegmental feature that conveys different meanings.
⭐ Key Takeaways
A student must remember that waveforms visually distinguish consonant manners of articulation (stop, fricative) and voicing (voicing bar). Vowel articulation is characterized by three factors: lip shape, tongue part raised, and tongue height. Acoustically, vowels are identified by their first two formant frequencies (F1 and F2), with F1 inversely related to height and F2 related to frontness. Length is a crucial phonemic feature in English, distinguishing short from long vowels, and diphthongs are glides between two vowel qualities. Suprasegmental features like pitch and stress operate beyond individual sounds, affecting meaning at the word and sentence level.
🧠 Quick Revision Questions
- How can you distinguish a voiced stop from a voiceless stop in a waveform?
- What are the three key articulatory parameters used to describe vowel sounds?
- Explain the relationship between the first formant (F1) and vowel height.
- Give one example of a phonemic long/short vowel contrast in English and state the two words that illustrate it.
- What is the difference between a segmental feature (like a consonant) and a suprasegmental feature (like intonation)?
📘 Lecture 7 — PHONEMIC AND PHONETIC TRANSCRIPTION-I
📖 Overview: This lecture introduces the fundamental concepts of phonetic and phonemic transcription, focusing on the International Phonetic Alphabet (IPA). It explains why transcription is essential in linguistics, describes the IPA chart, and provides detailed instruction on transcribing both vowels and consonants in English, specifically for the BBC accent.
🗂️ Topics Covered
The lecture begins by explaining the importance of transcription, distinguishing between phonemic and phonetic types. It then introduces the International Phonetic Association and its IPA chart, which is a standardized tool for transcribing speech sounds. Following this, the lecture provides detailed explanations and examples for transcribing English vowels (short, long, and diphthongs) and consonants (plosives, affricates, fricatives, nasals, and approximants).
📝 Lecture Summary
Topic-037: Why Transcribe?
Transcription is a crucial tool in phonetics and phonology. It is the process of writing down a spoken utterance using a specific set of symbols. There are two main types: phonemic (broad) transcription, which only uses symbols for the phonemes of a language, and phonetic (narrow) transcription, which uses a full range of symbols to capture fine phonetic detail. Transcription is essential for understanding sound-symbol correspondence and is a key technique for documenting and describing a language's phonemes.
Topic-038: Introduction to IPA
The International Phonetic Association (IPA) was established in 1886 by teachers and practitioners to improve spoken language teaching using phonetics. It had a revolutionary impact on language classrooms, shifting focus from written proficiency to spoken forms. The IPA remains a major international learned society, maintaining a research journal and a website with tools for phonetic study. It is best known for maintaining the International Phonetic Alphabet (IPA chart) , a specific set of alphabets used for transcription. 🔑 Definition — IPA: The International Phonetic Alphabet, a standardized system of symbols used to represent the sounds of spoken language.
Topic-039: Explaining IPA Chart
Since 1886, the IPA has continuously updated the IPA charts, with the last revision in 2015. The charts are used to transcribe not only segments (vowels and consonants) but also diacritics (detailed phonetic variation) and suprasegmental features (like stress and intonation). The lecture includes the IPA chart for consonants as a reference tool.
Topic-040: Transcription of Vowels
For transcribing vowels, the IPA vowel chart is used. The BBC accent of English is described as having short vowels, long vowels, and diphthongs.
- There are seven short vowels: /ɪ/ (pit), /e/ (pet), /æ/ (pat), /ʌ/ (putt), /ɒ/ (pot), /ʊ/ (put), /ə/ (another).
- There are five long vowels: /iː/ (bean), /ɑː/ (barn), /ɔː/ (born), /uː/ (boon), /ɜː/ (burn).
- There are eight diphthongs: /eɪ/ (bay), /aɪ/ (buy), /ɔɪ/ (boy), /əʊ/ (no), /aʊ/ (now), /ɪə/ (peer), /eə/ (pair), /ʊə/ (poor).
Topic-041: Transcription of Consonants
For consonant sounds in the BBC accent of English, there are 24 symbols. They are categorized as follows:
- Plosives (6): /p/ /b/ /t/ /d/ /k/ /g/
- Affricates (2): /ʧ/ /ʤ/
- Fricatives (9): /f/ /v/ /θ/ /ð/ /s/ /z/ /ʃ/ /ʒ/ /h/
- Nasals (3): /m/ /n/ /ŋ/
- Approximants (4): /l/ /r/ /w/ /j/
💡 Why this matters: Knowing these 24 symbols is essential for accurately representing every consonant sound in standard English pronunciation.
Topic-042: Transcription of Consonants: Explanation
Here are example words for each consonant category:
- Plosives: /p/ pin, /b/ bin, /t/ tin, /d/ din, /k/ kin, /g/ gum.
- Affricates: /ʧ/ chain, /ʤ/ Jane.
- Fricatives: /f/ fine, /v/ vine, /θ/ think, /ð/ this, /s/ seal, /z/ zeal, /ʃ/ sheep, /ʒ/ measure, /h/ how.
- Nasals: /m/ sum, /n/ sun, /ŋ/ sung.
- Approximants: /l/ light, /r/ right, /w/ wet, /j/ yet.
⭐ Key Takeaways
The most critical takeaway is the clear distinction between phonemic (broad) and phonetic (narrow) transcription. You must memorize the 44 phonemes of BBC English: 20 vowels (including seven short, five long, and eight diphthongs) and 24 consonants (categorized as plosives, affricates, fricatives, nasals, and approximants). The IPA chart is a living tool maintained since 1886 for transcribing all world languages. Practice transcribing words from each vowel and consonant category to build fluency. Finally, transcription is not just an academic exercise; it is a foundational skill for documenting languages, teaching pronunciation, and understanding sound-symbol correspondence.
🧠 Quick Revision Questions
- What is the fundamental difference between a phonemic (broad) and a phonetic (narrow) transcription?
- Name the three categories of vowel sounds in the BBC accent of English, and state how many sounds are in each category.
- List the 24 consonant symbols of BBC English, grouped by their manner of articulation.
- What does IPA stand for, and when was the International Phonetic Association founded?
- Provide an example word for each of the eight diphthongs in English.
📘 Lecture 8 — PHONEMIC AND PHONETIC TRANSCRIPTION-II
📖 Overview: This lecture expands on the distinction between broad (phonemic) and narrow (phonetic) transcription, introducing the IPA’s official resource story, "The North Wind and the Sun," as a standard text for transcription practice. Students learn to apply both transcription types to the same text, using specific symbols for the BBC accent of English and understanding the purpose of diacritics in narrow transcription.
🗂️ Topics Covered
This lecture covers the definitions and differences between broad and narrow transcription, introduces the IPA resource "The North Wind and the Sun" story with its explanation, provides practice in phonemic transcription using BBC English symbols, and concludes with principles and an example of phonetic (narrow) transcription for the same text.
📝 Lecture Summary
Topic-043: Broad and Narrow Transcription
Two main kinds of transcription are recognized: broad (phonemic) and narrow (phonetic). Conventionally, square brackets [ ] enclose phonetic transcription, while oblique lines / / enclose phonemic transcription. In broad transcription, sounds are symbolized based on their linguistic functions in a language, without detailing the physical features of an individual sound. Phonemic transcription represents only the units that account for differences in meaning, e.g., /pin/, /pen/, /pæn/. In narrow transcription, sounds are symbolized based on their articulatory/auditory identity, regardless of their function in a language (sometimes called impressionistic transcription). The aim here is to identify sounds as such, focusing on phonetic variation.
🔑 Definition — Broad Transcription (Phonemic): A transcription that symbolizes sounds based on their linguistic function, using slashes / /, and representing only phonemes that change meaning. 📐 Formula: Broad /k/ vs. Narrow [k] → / / encloses phonemic; [ ] encloses phonetic. 📌 Example: The words /pin/, /pen/, /pæn/ are broad transcriptions showing phonemic differences (different vowels change meaning). Narrow transcription would add diacritics to show fine phonetic detail.
Topic-044: IPA Resource: The North Wind and the Sun Story
The homepage of the International Phonetics Association (The IPA) provides many helpful links, including sound files, fonts, and the IPA journal. The IPA offers various tools for the phonetic study of human languages, one of which is ‘The North Wind and the Sun’ story. To create a uniform system for describing language sounds, this text is recommended for transcription (both narrow and broad), especially for publishing IPA illustrations of languages.
Topic-045: IPA Resource: Explanation
The story ‘The North Wind and the Sun’ is provided for transcription practice. The full text is: The north wind and the sun were disputing which was the stronger when a traveler came along wrapped in a warm cloak. They agreed that the one who first succeeded in making the traveler take his cloak off should be considered stronger than the other. Then the north wind blew as hard as he could, but the more he blew the more closely did the traveler fold his cloak around him and at last the north wind gave up the attempt. Then the sun shined out warmly, and immediately the traveler took off his cloak. And so the north wind was obliged to confess that the sun was the stronger of the two.
Topic-046: IPA Story Practice
For transcription, use the symbols of the BBC accent of English. The vowels are: ɪ e æ ʌ ɒ ʊ ǝ; iː ɑː uː ɜː; eɪ aɪ ɔɪ ǝʊ aʊ ɪǝ eǝ ʊǝ. The consonants are: p b t d k g; f v θ ð s z ʃ ʒ h; m n ŋ l r w j ʧ ʤ.
Topic-047: IPA Story: Broad Transcription
To phonemically transcribe the story, use oblique lines/slashes / / and transcribe broadly without phonetic detail. The first sentence is transcribed as: /ðə norθ wɪnd ən ðə sʌn wər dɪspjutɪŋ/. Key considerations are: (1) using slashes to enclose transcription, and (2) transcribing only phonemic features. 💡 Why this matters: This is a foundational skill for representing the functional sound system of English.
Topic-048: IPA Story-Phonetic Transcription
This topic covers narrow or detailed transcription (phonetic transcription), where sounds are symbolized based on articulatory/auditory features, using square brackets [ ] and diacritics to capture phonetic variation. The first sentence is transcribed as: [ðə ˈnɔɹθ ˌwɪnd ən ə ˈsʌn wɚ dɪˈspjuɾɪŋ]. Key differences from broad transcription include the use of stress marks (ˈ and ˌ), diacritics (e.g., the flap [ɾ]), and the rhotic vowel [ɚ].
🔑 Definition — Narrow Transcription (Phonetic): A transcription that symbolizes sounds based on their articulatory/auditory features, using square brackets [ ] and diacritics from the IPA to show fine phonetic detail. 📌 Example: Broad /ðə norθ wɪnd ən ðə sʌn wər dɪspjutɪŋ/ vs. Narrow [ðə ˈnɔɹθ ˌwɪnd ən ə ˈsʌn wɚ dɪˈspjuɾɪŋ]. The narrow version shows stress patterns and a flap [ɾ] for the /t/ sound.
⭐ Key Takeaways
A student must remember that broad transcription uses slashes / / and represents only phonemes (meaning-distinguishing sounds), while narrow transcription uses brackets [ ] and includes phonetic detail like stress, diacritics, and allophones. The IPA provides "The North Wind and the Sun" story as a standard resource for practicing both types of transcription in any language. For BBC English, specific vowel and consonant symbols are used, and narrow transcription can show features like the flap [ɾ] or rhotic vowels. This distinction is crucial for accurately representing both the functional and physical aspects of speech sounds.
🧠 Quick Revision Questions
- What are the two main types of transcription, and what symbols enclose each?
- What is the primary difference in purpose between broad and narrow transcription?
- What is the name of the standard story provided by the IPA for transcription practice?
- In the broad transcription example /ðə norθ wɪnd ən ðə sʌn wər dɪspjutɪŋ/, what does the symbol /ð/ represent?
- How does the narrow transcription [ðə ˈnɔɹθ ˌwɪnd ən ə ˈsʌn wɚ dɪˈspjuɾɪŋ] differ from its broad counterpart (specifically for the word "disputing")?
📘 Lecture 09 — PHONEMIC AND PHONETIC TRANSCRIPTION-III
📖 Overview: This lecture covers the transcription of suprasegmental features beyond individual words, focusing on connected speech. It explains how stress, accent, phrase-level changes, and rhythm are transcribed and analyzed in continuous utterances, which is essential for accurate phonetic representation of normal speech.
🗂️ Topics Covered
This lecture covers transcription beyond words, including word stress with primary and secondary stress symbols, accent variation across dialects, phrase-level changes such as assimilation and elision, and the rhythmic patterns of connected speech exemplified through poetry. It emphasizes the need to account for pronunciation changes that occur when words are used in natural, flowing utterances.
📝 Lecture Summary
Topic-049: Transcription Beyond Words
In connected speech (normal utterances and conversations), important changes happen to sound units that are not present at the word level. For example, "and" becomes /n/ in phrases like "boys and girls," and /n/ becomes /m/ in "green bus." The features of connected speech include assimilation, rhythm, stress, elision, linking, tone, and intonation.
Topic-050: Transcription Beyond Words: Word Stress
Stress refers to the degree of force used in producing a syllable. The distinction is between stressed and unstressed syllables, with stressed syllables being more prominent due to increased loudness, length, and often pitch. The IPA symbol for primary stress is a raised vertical line [ˈ] placed before the stressed syllable (e.g., /ˈsnɪp.ɪt/, /ɪgˈzɪst/, /prənʌnsiˈeiʃən/). The symbol for secondary stress is a lowered vertical line [ˌ] placed before the stressed syllable (e.g., /ˌmɪnɪmaɪˈzeɪʃən/).
🔑 Definition — Stress: The degree of force used in producing a syllable; stressed syllables are more prominent than unstressed ones. 📐 Formula: Primary stress = [ˈ] before syllable; Secondary stress = [ˌ] before syllable → These symbols mark which syllable receives the most or second-most force in a word. 📌 Example: In /prənʌnsiˈeiʃən/, the primary stress falls on the syllable "ei" (marked with [ˈ]), while the other syllables are unstressed or have secondary stress.
Topic-051: Transcription Beyond Words: Accent
Accent refers to the typical way of pronunciation that identifies a community, such as BBC English, American English, or Chinese English. A single language can have many possible transcriptions based on accent differences. Four types of variation among accents must be recognized: a. Difference in phonological inventories: e.g., /strʌt/ vs. /strʊt/ (vowel differences). b. Difference in phonetic features: e.g., pronouncing /t/ as [ʔ] (glottal stop). c. Phonological distribution: e.g., rhotic vs. non-rhotic accents (presence or absence of /r/ after vowels). d. Lexical distribution: e.g., /θ/ and /ð/ differences (England and Wales use /ð/, while Scottish accents use /θ/ in certain words).
💡 Why this matters: Accurate transcription must account for accent-specific variations to reflect the actual pronunciation of a speaker or community.
Topic-052: Transcription Beyond Words: Phrases
Words in phrases undergo transitions and changes due to contact. For example, in "ten green bottles," the /n/ in "ten" changes to /ŋ/ in anticipation of the /g/ in "green" (becoming [teŋ]), and the /n/ in "green" changes to /m/ in anticipation of the /b/ in "bottles" (becoming [griːm]). This process is called assimilation — sounds become more similar to neighboring sounds to simplify pronunciation. Other changes include elision (omission of sounds), linking (connecting sounds across word boundaries), epenthesis (insertion of sounds), and liaison (linking of final and initial sounds).
🔑 Definition — Assimilation: A process where a sound changes to become more similar to a neighboring sound, reflecting an economy of effort in connected speech. 📌 Example: In "ten green bottles," /ten/ + /griːn/ → [teŋ griːm bɒtəlz]; the /n/ changes to /ŋ/ before /g/, and /n/ changes to /m/ before /b/.
Topic-053: Transcription Beyond Words: Rhythm and Beyond
Connected speech has noticeable events at regular intervals, known as rhythmicality. These regularities involve patterns of stressed vs. unstressed syllables, syllable length (long vs. short), or pitch (high vs. low), or combinations thereof. The most regular patterns, found in poetry, are called metrical. The poem "Ten green bottles" illustrates how rhythmic features must be considered when transcribing connected speech:
- "Ten green bottles / Hanging on the wall / And if one green bottle / Should accidentally fall / There’d be nine green bottles / Hanging on the wall" This requires analyzing stress patterns, syllable timing, and pitch changes that occur in natural, flowing speech.
🔑 Definition — Rhythm: The pattern of stressed and unstressed syllables or long and short syllables occurring at regular intervals in speech. 📐 Example: In poetry, metrical patterns create a regular rhythm; in "Ten green bottles," the stress falls on "Ten," "green," and "bottles" in a predictable pattern.
⭐ Key Takeaways
The most critical concepts from this lecture are: connected speech requires transcribing suprasegmental features beyond word-level segments, including stress, accent, and phrase-level changes. Stress is marked with primary [ˈ] and secondary [ˌ] symbols before the stressed syllable. Accent variations across dialects affect phonological inventories, phonetic features, distribution, and lexical choices, requiring careful transcription adjustments. Phrase-level changes like assimilation (e.g., /n/ becoming /m/ in "green bus") reflect simplification processes. Rhythm, including metrical patterns in poetry, must be captured to accurately represent the natural flow of continuous utterances.
🧠 Quick Revision Questions
- What are the main features of connected speech that must be transcribed beyond word level?
- How do you mark primary and secondary stress in IPA transcription? Provide an example word.
- What are the four types of accent variation among different dialects of a language?
- Explain the process of assimilation using the phrase "ten green bottles" as an example.
- What is meant by "rhythm" in connected speech, and how does it relate to metrical patterns in poetry?
📘 Lecture 10 — The Consonants of English-I
📖 Overview: This lecture introduces the 24 consonants of English (RP accent), explaining how they are described through three key features: voicing, manner of articulation (MoA), and place of articulation (PoA). It systematically covers stop consonants, fricatives, affricates, nasals, and approximants, providing a foundational understanding for phonetic analysis and transcription.
🗂️ Topics Covered
The lecture begins by presenting a comprehensive table of English consonants categorized by voicing, place, and manner of articulation. It then delves into the specific features of stop consonants (including oral and nasal stops), followed by an exploration of fricatives (with fortis/lenis distinctions and the concept of obstruents). The section on affricates explains their dual plosive+fricative nature, while nasals are described through their articulatory mechanism. Finally, approximants are discussed, covering semivowels and liquids.
📝 Lecture Summary
Topic-054: The Consonants of English
The RP (Received Pronunciation) accent of English contains 24 consonants. These are described using three primary features: voicing (whether the vocal cords vibrate), manner of articulation (how the airflow is obstructed), and place of articulation (where in the vocal tract the obstruction occurs). A standard reference table shows these consonants arranged by place (e.g., bilabial, labiodental, dental, alveolar, postalveolar, palatal, velar, glottal) and manner (plosive, fricative, affricate, nasal, approximant, lateral). In such tables, consonants on the left side of each cell are typically voiceless, while those on the right are voiced.
Topic-055: Stop Consonants
A stop is a consonant made with a complete closure in the oral cavity, blocking airflow. While often synonymous with plosive, some phoneticians use "stop" more broadly to include nasal stops (where air escapes through the nose). English has nine stops: six oral plosives and three nasal stops.
- Oral plosives occur in three places: bilabial (/p/, /b/), alveolar (/t/, /d/), and velar (/k/, /g/).
- Nasal stops are also found in three places: bilabial (/m/), alveolar (/n/), and velar (/ŋ/).
- The glottal stop /ʔ/ also appears in some English varieties, e.g., in 'beaten' [ˈbɪʔn̩].
- English voiceless stops (/p, t, k/) are aspirated (produced with a puff of air) when they occur at the beginning of a stressed syllable, e.g., in "pie" [pʰaɪ], "tie" [tʰaɪ], "kie" [kʰaɪ].
🔑 Definition — Nasal stop: A stop consonant where air is released through the nose due to a lowered velum, while a complete closure is made in the oral cavity. 📐 Table: Oral stops: /p, b, t, d, k, g/; Nasal stops: /m, n, ŋ/. 📌 Example: The word "pat" has an aspirated voiceless bilabial stop at the beginning: [pʰæt]. The word "mat" begins with a voiced bilabial nasal stop: [mæt].
Topic-056: Fricatives
A fricative is a sound produced when two articulators are brought so close together that the airflow becomes turbulent, creating audible friction. There is no complete closure, only a narrow constriction. English has several fricatives, both voiced and voiceless.
- Common fricatives and their examples: /f/ (fin), /v/ (van), /θ/ (thin), /ð/ (this), /s/ (sin), /z/ (zoo), /ʃ/ (ship), /ʒ/ (measure), /h/ (hoop).
- The fricative manner of articulation produces a wider range of speech sounds than any other.
- Fricatives are divided into fortis (strong, voiceless: /f, s, θ, ʃ, h/) and lenis (weak, voiced: /v, z, ð, ʒ/) based on articulatory energy.
- Obstruents is a super-category that includes stops and fricatives. They share three key properties: (1) they shorten preceding vowels (vowels are shorter before voiceless obstruents), (2) voiceless obstruents at the end of a syllable are longer than their voiced counterparts (e.g., race vs. rays), and (3) they are voiced only if adjacent sounds are also voiced (e.g., dogs [dɒɡz]).
🔑 Definition — Fortis: A term for voiceless fricatives, which are produced with greater muscular energy and airflow. Lenis: A term for voiced fricatives, produced with less muscular energy. 📌 Example: In the word "sip" /sɪp/, /s/ is a voiceless fortis fricative. In "zip" /zɪp/, /z/ is a voiced lenis fricative.
Topic-057: Affricates
An affricate is a complex consonant that begins as a plosive (complete closure) and is released as a fricative (narrow constriction) at the same place of articulation. It is a single phonetic segment, despite its two-step production.
- English has two affricates: the voiceless /tʃ/ (as in the beginning and end of church /tʃɜ:tʃ/) and the voiced /dʒ/ (as in the beginning and end of judge /dʒʌdʒ/).
- Both affricates are classified as post-alveolar in place of articulation.
🔑 Definition — Affricate: A single consonant sound that consists of a plosive immediately followed by a fricative, both at the same place of articulation. 📌 Example: The word "cheap" begins with the voiceless affricate /tʃ/. The word "jeep" begins with the voiced affricate /dʒ/.
💡 Why this matters: Affricates are often confused with sequences of two separate sounds (e.g., /t/ + /ʃ/ in "hits"), but they function as a single phoneme, changing word meanings.
Topic-058: Nasals
Nasals are consonant sounds where the velum (soft palate) is lowered, allowing air to escape through the nose. Two articulatory actions are required: (1) the velum is lowered for nasal airflow, and (2) a complete closure is made somewhere in the oral cavity to prevent oral escape.
- English has three voiced nasal stops: bilabial /m/, alveolar /n/, and velar /ŋ/.
- All are voiced sounds.
🔑 Definition — Nasal: A consonant sound produced with a lowered velum, allowing air to flow through the nose, while the oral cavity is blocked. 📌 Example: The word "sing" /sɪŋ/ ends with the velar nasal /ŋ/. The word "sum" /sʌm/ ends with the bilabial nasal /m/.
Topic-059: Approximants
Approximants are consonants that create very little obstruction to the airflow, so they do not produce friction or turbulent noise. They are traditionally divided into two groups:
- Semivowels: Sounds like /w/ (as in wet) and /j/ (as in yet), which are phonetically very similar to close vowels /u/ and /i/ but function as consonants, produced as a rapid glide.
- Liquids: Sounds like the lateral /l/ (as in lead) and the /r/ sound (as in read). These have a more identifiable constriction but still do not obstruct enough to create fricative noise.
The BBC accent has four approximants:
- Bilabial: /w/ (whack)
- Alveolar: /l/ (lack) and /r/ (rack)
- Palatal: /j/ (yak)
Sometimes, experts differentiate among various kinds of /r/ (tap, flap, trill).
🔑 Definition — Approximant: A consonant where the articulators are close but not narrow enough to cause turbulent airflow; the sound is made with minimal obstruction. 📌 Example: The word "yes" begins with the palatal approximant /j/. The word "well" begins with the bilabial approximant /w/.
⭐ Key Takeaways
You must memorize that English has 24 consonants, all described by voicing, manner, and place of articulation. The nine stops include six oral plosives (/p, b, t, d, k, g/) and three nasal stops (/m, n, ŋ/), with voiceless stops aspirated word-initially. Fricatives are divided into fortis and lenis, and together with stops, form the class of obstruents which affect vowel length and final consonant length. The two affricates (/tʃ/ and /dʒ/) are single segments that combine a plosive and fricative at the post-alveolar place. Finally, the four approximants (/w, j, l, r/) are distinguished by minimal airflow obstruction, with semivowels (/w, j/) acting like gliding vowels and liquids (/l, r/) having more notable constriction.
🧠 Quick Revision Questions
- How many consonants are there in the RP accent of English, and what three features are used to describe them?
- What is the difference between an oral stop and a nasal stop? Give one example of each from English.
- What is the key distinction between a fricative and an affricate?
- Name the four approximants in the BBC accent of English and list their places of articulation.
- Explain what an obstruent is, and list three properties shared by stops and fricatives.
📘 Lecture 11 — The Consonants of English-II
📖 Overview: This lecture explores how speech sounds interact in connected speech through overlapping articulatory gestures and co-articulation. It introduces the key rules governing English consonant allophones, explaining why the same phoneme can sound different depending on its phonetic context. Understanding these rules is essential for accurate phonetic transcription and phonetic analysis.
🗂️ Topics Covered
The lecture begins with overlapping gestures in speech production, where articulatory movements for adjacent sounds blend together. It then discusses aspects of connected speech that create allophonic variation, followed by the specific phenomenon of co-articulation. The majority of the lecture is dedicated to presenting and explaining 19 formal rules for English consonant allophones, covering aspiration, voicing, length, glottalization, syllabic consonants, and place assimilation.
📝 Lecture Summary
Topic-060: Overlapping Gestures
Speech sounds are produced with movements of the articulators, and sounds are often described in terms of their articulatory gestures. Sounds are not static; they are movements. This idea makes it easier to understand the overlapping of sounds in terms of their articulatory gestures. Try saying words twice, dwindle, quick and analyze the rounding of your lips for sound /w/. In each of these three words, the first stop sounds are slightly rounded (when they are clustered with /w/ — /tw/, /dw/ and /kw/ respectively). In these words, there is a tendency for gestures to overlap with those for adjacent sounds (stops with bilabial /w/ in this case). This kind of gestural overlapping, in which a second gesture starts during the first gesture, is sometimes also called anticipatory co-articulation. The articulatory gesture for the approximant sound is anticipated during the articulatory gesture for the stop. The same kind of anticipatory overlapping takes place in words like tree and dream (compare them with tea and deem). In phonology, overlapping refers to the possibility when a phone may be assigned to more than one phoneme (phonemic overlapping). As a notion, overlapping was introduced by American structural linguists in the 1940s.
🔑 Definition — Overlapping (gestural): A phenomenon in which the articulatory gesture for one sound starts during the gesture for an adjacent sound, often called anticipatory co-articulation. 📌 Example: In the word twice /twaɪs/, the /t/ is slightly rounded because the lip rounding gesture for /w/ overlaps with the /t/.
Topic-061: Aspects of Connected Speech
Overlapping is a common feature of connected speech. In a rapid (connected) speech, overlapping between sounds results in the positions of some parts of the vocal tract being influenced quite a lot by neighboring targets, thus creating various forms (allophones) for one phoneme. Keeping in mind this possibility of overlapping, a phoneme is an abstract unit that may be realized in several different ways (forms — allophones). Similarly, the differences between various allophones of a phoneme can be explained in terms of targets and overlapping gestures. The difference between two different forms of /k/ sound (as the [k] in key and the [k] in caw) may be simply due to their overlapping with different vowels in context. Similarly, the alveolar [n] in ten is different than the dental [n̪] in tenth. Both are the result of aiming at the same target, but in tenth, the realization of the phoneme /n/ is influenced by the dental target required for the following sound.
🔑 Definition — Connected speech: Natural, rapid speech in which sounds are not produced in isolation but are influenced by neighboring sounds, leading to overlapping and allophonic variation. 📌 Example: The /n/ in ten [tɛn] is alveolar, but the /n/ in tenth [tɛn̪θ] is dental because the tongue tip anticipates the dental target for the following /θ/.
Topic-062: Co-articulation
An articulation is an articulatory phenomenon which involves a simultaneous overlapping of more than one point in the vocal tract, as in the co-ordinate stops (/pk/, /bg/, /pt/ and /bd/) often heard in some languages from West Africa. Co-articulation, at times, leads to creating a difference between two allophones (which is actually the result of aiming at different targets). In experimental phonetics, co-articulation is a way of finding out how the brain controls the production of speech sounds. When we speak, many muscles are active at the same time and sometimes the brain tries to make them do things at a time that they are not capable of. For example, in the word mum /mʌm/, the vowel phoneme is one that is normally pronounced with the soft palate (velum) raised to prevent the escape of air through the nose, while the two m phonemes must have the soft palate lowered. Thus, the soft palate cannot be possibly raised so quickly, and, as a result, the vowel is most likely to be pronounced with the soft palate (velum) still lowered — making the vowel a nasalised one. Thus, in this case, the nasalization is a co-articulation effect which is caused by the nasal consonants in context (environment). Another example is the lip-rounding as discussed in Topic 60 above.
🔑 Definition — Co-articulation: A simultaneous overlapping of articulatory gestures from more than one point in the vocal tract, often resulting in allophonic variation (e.g., nasalization of a vowel between nasal consonants). 💡 Why this matters: Co-articulation demonstrates that speech planning involves coordinating multiple articulators simultaneously, and the physical limitations of the vocal tract create predictable allophonic patterns. 📌 Example: In mum /mʌm/, the velum cannot raise quickly enough between the two /m/s, so the vowel /ʌ/ becomes nasalized [ʌ̃].
Topic-063: Rules for English Consonant Allophones (ECA)
Based on the above discussion on overlapping and co-articulatory gestures, the rules for English consonantal allophones are summarized here. Remember that it is just a list of a set of formal statements simply describing the behavior of a language. These are not the kind of prescriptive grammar rules that people are expected to abide by.
- Consonants are longer when at the end of a phrase (e.g., bib, did, don and nod).
- Voiceless stops (e.g., p, t, k) are aspirated when they are syllable initial (pip, test, kick).
- Voiced obstruents (b, d, g, v, ð, z, ʒ) are voiced only when they occur at the end of an utterance or before a voiceless sound.
- Voiced stops (b, d, g) and affricate (dʒ) are voiceless when they are syllable initial (except when immediately preceded by a voiced sound — compare a day with this day).
- Voiceless stops (p, t, k) are unaspirated after /s/ in words such as spew, stew and skew.
- Voiceless obstruents (p, t, k, tʃ, f, θ, s, ʃ) are longer than their voiced counterparts (b, d, g, dʒ, v, ð, z, ʒ) at the end of a syllable (e.g., cap — cab and back — bag).
- Approximants (w, r, j, l) are at least partially voiceless when they occur after initial voiceless stop sounds (e.g., play, twin, cue).
- The gestures for consecutive stops overlap, so that stops are unexploded when they occur before another stop (e.g., apt and rubbed).
- In many accents of English, syllable final voiceless stops /p, t, k/ are accompanied by an overlapping glottal stop gesture (e.g., tip, pit, kick).
- /t/ is replaced by a glottal stop when it occurs before an alveolar nasal (e.g., beaten).
- Nasals are syllabic at the end of a word — after an obstruent (e.g., leaden, chasm).
- The lateral /l/ is syllabic at the end of a word — a consonant (e.g., paddle, whistle).
🔑 Definition — Aspiration: A burst of air following the release of a voiceless stop, transcribed with a superscript h [pʰ, tʰ, kʰ]. 📐 Rule 2: Voiceless stops /p, t, k/ → [pʰ, tʰ, kʰ] / #__V (syllable-initial, before a vowel). Plain-English meaning: When a voiceless stop starts a syllable, it is accompanied by a puff of air. 📌 Example: pip [pʰɪp] — the initial /p/ is aspirated, but the final /p/ is not. 📐 Rule 5: Voiceless stops /p, t, k/ → [p, t, k] (unaspirated) / /s/ __. Plain-English meaning: After /s/, voiceless stops lose their aspiration. 📌 Example: spew [spju] — the /p/ after /s/ is unaspirated.
📐 Rule 11: Nasals → syllabic / C__# (after an obstruent, at word end). Plain-English meaning: When a nasal follows an obstruent at the end of a word, it becomes the nucleus of a syllable. 📌 Example: leaden /lɛdən/ — the final /n/ becomes syllabic [n̩], so the phonetic transcription is [lɛdn̩].
Topic-064: Rules for English Consonant Allophones (ECA): Explanation
- An alveolar stop becomes a voiced tap when it occurs between two vowels, the second of which is unstressed (winter — winner).
- An alveolar consonant becomes dental before a dental consonant (eighth, tenth, wealth).
- Alveolar stops are reduced or omitted when between two consonants (/moʊst pIpl/ — /moʊs pIpl/).
- A homorganic voiceless stop may occur after a nasal before a voiceless fricative followed by an unstressed vowel in the same word (e.g., hearing /t/ in both agency and grievances).
- A consonant is shortened when it is before an identical consonant (e.g., /k/ in cap and kept).
- Velar stops become more frontal before more frontal vowels (e.g., clap and talc).
- The lateral /l/ is velarized after a vowel or before a consonant at the end of a word.
🔑 Definition — Tap: A quick flap of the tongue tip against the alveolar ridge, represented as [ɾ]. 📐 Rule 1: Alveolar stops /t, d/ → [ɾ] / V__V (where the second vowel is unstressed). Plain-English meaning: /t/ and /d/ become a tap when they are between two vowels and the second vowel is unstressed. 📌 Example: winter /wɪntər/ → [wɪnɾɚ] (can be homophonous with winner [wɪnɚ]).
🔑 Definition — Velarized: A secondary articulation in which the back of the tongue is raised toward the velum while producing another sound, symbolized as [ɫ]. 📐 Rule 7: /l/ → [ɫ] / V__ or __C# (after a vowel or before a consonant at word end). Plain-English meaning: At the end of a syllable, /l/ has a "dark" sound with the back of the tongue raised. 📌 Example: talc [tʰælk] — the /l/ is velarized, whereas it is "clear" in clap [klæp].
⭐ Key Takeaways
The most critical concepts from this lecture are: (1) Overlapping gestures and co-articulation explain why phonemes have different allophones — sounds are not produced in isolation but are influenced by neighboring articulatory movements. (2) The 19 rules for English consonant allophones describe predictable patterns, such as aspiration of voiceless stops syllable-initially (Rule 2), devoicing and unaspiration after /s/ (Rule 5), and the tap rule for /t/ and /d/ between vowels (Rule 064-1). (3) Co-articulation effects like vowel nasalization in mum and lip-rounding in clusters like /tw/ demonstrate how the brain plans speech and how the vocal tract's physical constraints create allophonic variation. (4) Syllabic nasals and laterals (Rules 11-12) and glottal replacement (Rules 9-10) are common in connected speech and must be recognized in transcription. (5) Remember that these rules are descriptive, not prescriptive — they describe how English is naturally spoken, not how it "should" be spoken.
🧠 Quick Revision Questions
- What is the difference between overlapping gestures and co-articulation? Provide an example of each from the lecture.
- Under what conditions are voiceless stops /p, t, k/ aspirated, and when are they unaspirated?
- Explain the "tap rule" (Rule 064-1). Why does winter sound like winner in many accents of English?
- Describe three different contexts in which /t/ can be realized as a glottal stop [ʔ] according to the lecture's rules.
- What does it mean for a nasal or lateral to be "syllabic"? Provide one example for each from Rules 11 and 12.
📘 Lecture 12 — THE CONSONANTS OF ENGLISH-III
📖 Overview: This lecture introduces diacritics—small marks added to phonetic symbols for detailed transcription. It explains key diacritics for aspiration, nasalization, and velarization, and their importance in narrow transcription, distinguishing between phonemic and allophonic features in English and other languages.
🗂️ Topics Covered
The lecture covers an introduction to diacritics and their role in detailed (narrow) transcription, with six important diacritic symbols and examples. It then focuses on three specific diacritic-related processes: aspiration (a puff of air after plosive release), nasalization (nasal influence on adjacent sounds via coarticulation), and velarization (secondary articulation with tongue back raising), explaining their articulation, phonetic transcription, and cross-linguistic significance.
📝 Lecture Summary
Topic-065: Introduction to Diacritics
A diacritic is a small mark added to a phonetic symbol to show the way it is spoken in accurate and detailed transcription. Diacritics include accent marks (e.g., ́ ` ^), the sign of devoicing [o] (a small circle below), and the sign of nasalization [~] (a tilde above). These marks may be placed over a symbol, under it, before it, after it, or through it. The International Phonetic Association (IPA) recognizes a wide range of such marks for both vowels and consonants. For vowels, diacritics indicate differences in frontness, backness, closeness or openness, lip-rounding or unrounding, nasalization, and centralization. For consonants, diacritics are used for voicing or voicelessness, advanced or retracted place of articulation, aspiration, and many other aspects. These small marks are essential for narrow transcription, which provides a more detailed phonetic representation than broad transcription.
Topic-066: Diacritics and Detailed Transcription
For a detailed transcription, diacritics are used to narrow the meaning of a symbol. The lecture lists six important diacritics for detailed transcription exercises.
🔑 Definition — Diacritic: A small mark added to a phonetic symbol to indicate a specific phonetic feature (e.g., voicelessness, aspiration, dental articulation, nasalization, velarization, or syllabicity).
📐 Six Key Diacritics Table:
| S. No. | Feature | Symbol | Description | Example Transcription |
|---|---|---|---|---|
| 1 | Voiceless | ̥ | Small circle below | quick /kw̥ɪk/ |
| 2 | Aspirated | ʰ | Small /h/ above | kiss /kʰɪs/ |
| 3 | Dental | ̪ | Dental sign below | health /həl̪θ/ |
| 4 | Nasalized | ̃ | Tilde symbol above | man /mæ̃n/ |
| 5 | Velarized | ̴ | Tilde symbol through | pill /pʰɪl̴/ |
| 6 | Syllabic n | ̩ | Small vertical line below | mitten /mɪʔn̩ɪ/ |
📌 Example: In the word "mitten" /mɪʔn̩ɪ/, the syllabic n diacritic [ ̩] indicates that the /n/ acts as a syllable nucleus (like a vowel), without a vowel sound between the /n/ and the following /ɪ/. This is narrow transcription because it shows the exact pronunciation, not just the phonemic representation.
Topic-067: Aspiration
Aspiration is a puff of noise made when a consonantal constriction is released and air is allowed to escape relatively freely. In English, the plosives /p t k/ are aspirated at the beginning of a syllable (e.g., "kiss" /kʰɪs/). Phonetically, aspiration is the result of the vocal cords being widely parted at the time of the articulatory release. Aspiration is allophonic in English (it does not distinguish meaning), while in languages like Urdu, it is phonemic (it can change the meaning of a word). Pronunciation teachers used to practice aspirated plosives by asking learners to blow out a candle flame with the rush of air after /p t k/. A different articulation is used for voiced aspirated plosives found in many Indian languages (often spelled as ‘bh’, ‘dh’, ‘gh’). In these sounds, after the release of the constriction, the vocal folds vibrate to produce voicing but are not firmly pressed together, allowing a large amount of air to escape at the same time, producing a “breathy” quality. It is not only stops that are aspirated; both unaspirated and aspirated affricates also exist in Urdu.
🔑 Definition — Aspiration: A puff of air released after the release of a consonantal constriction, especially with voiceless plosives /p t k/ in English, represented by the diacritic [ʰ] above the symbol.
💡 Why this matters: English speakers distinguish aspirated from unaspirated plosives allophonically (e.g., "top" vs. "stop"), but in other languages, this difference is phonemic, so accurate transcription is crucial for understanding cross-linguistic phonetic contrasts.
Topic-068: Nasalization
Nasalization is an articulatory process whereby a sound is made ‘nasal’ when the air is passing through the nasal cavity due to the influence of an adjacent nasal sound. This is an example of anticipatory coarticulation, where the soft palate is lowered in anticipation of a following nasal consonant. For example, in the word "man" /mæ̃n/, the vowel /æ/ may be articulated with the soft palate lowered throughout, making it a nasalized vowel. The lecture emphasizes a distinction between a nasal sound and a nasalized sound. A sound is nasalized when the nasality comes from other sounds (as in the vowel above), whereas the term nasal suggests that the nasality is an essential identifying feature of the sound (e.g., nasal consonants /m n ŋ/ in English). In Urdu, there are many nasal sounds. A nasalized consonant is a consonant which, though normally oral, is articulated in a nasal manner because of some adjacent (nasal) sound.
🔑 Definition — Nasalization: The process by which a normally oral sound (especially a vowel) becomes nasal due to the influence of an adjacent nasal consonant, represented by the tilde diacritic [̃] above the symbol.
📌 Example: In the word "hand" /hænd/, the vowel /æ/ is nasalized before the nasal consonant /n/, though it may not be transcribed in broad transcription. In narrow transcription, it would be written as /hæ̃nd/.
Topic-069: Velarisation
Velarisation is a co-articulation process whereby a constriction in the vocal tract is added to the primary constriction which gives a consonant its place of articulation. More specifically, velarisation is an example of secondary articulation. In the case of English “dark /l/” [l̴], the /l/ phoneme is produced with its usual primary constriction in the alveolar region, but the back of the tongue is raised for an /u/ vowel sound, creating a secondary constriction. This is transcribed with the velarized diacritic [ ̴] (a tilde through the symbol). Examples include the contrast between "life" /laɪf/ (with clear /l/, no velarization) and "file" /faɪl̴/ (with dark /l/, velarized), and "clap" /klæp/ (clear) vs. "talc" /tæl̴k/ (dark). Velarisation is a very common feature of Arabic and is important and interesting for acoustic analysis.
🔑 Definition — Velarisation: A secondary articulation process where the back of the tongue is raised towards the velum (soft palate) while producing a primary articulation (e.g., alveolar /l/), creating a "dark" quality, represented by the tilde-through diacritic [ ̴].
📌 Example: In English, the /l/ in "leaf" (clear /l/) is not velarized, but the /l/ in "feel" (dark /l/) is velarized. In narrow transcription, "feel" would be transcribed as /fiːl̴/.
⭐ Key Takeaways
This lecture emphasizes that diacritics are essential for narrow phonetic transcription, allowing linguists to capture subtle but important phonetic details that broad transcription misses. The six key diacritics—voiceless, aspirated, dental, nasalized, velarized, and syllabic n—must be memorized and applied correctly. Aspiration, nasalization, and velarization are all co-articulatory processes that involve secondary articulatory gestures; they are allophonic in English but may be phonemic in other languages. Understanding the difference between a "nasal" sound (inherently nasal, like /m n ŋ/) and a "nasalized" sound (a normally oral sound made nasal by context, like a vowel before a nasal consonant) is critical. Finally, velarization explains the "dark /l/" phenomenon in English, which is a common feature that distinguishes varieties of English and appears in other languages like Arabic.
🧠 Quick Revision Questions
- What is a diacritic, and why are diacritics important for narrow transcription?
- What is the difference between an aspirated and an unaspirated plosive? Give an English example.
- In the word "man", why is the vowel /æ/ often nasalized? What is this process called?
- What is velarisation, and how does it distinguish the /l/ in "leaf" from the /l/ in "feel"?
- Name the six diacritics introduced in Topic-066 and give one example word for each.
📘 Lecture 13 — English Vowels-I
📖 Overview: This lecture introduces the fundamental features of English vowels, focusing on the challenges of describing vowel quality and the continuous nature of vowel space. It explains why vowels are difficult to transcribe across different accents and explores the key auditory dimensions used to characterize them. This matters because understanding vowel systems is essential for accurate phonetic transcription and for recognizing differences between varieties of English.
🗂️ Topics Covered
This lecture covers an introduction to English vowels, including the discrepancy in vowel counts across varieties and the importance of vowel length and quality. It then explains the concept of vowel quality as an auditory feature and the problem of precisely describing tongue position. The auditory vowel space is introduced, using cardinal vowels as reference points. Finally, the lecture compares American and British English vowels, noting specific differences in production and perception.
📝 Lecture Summary
Topic-070: Introduction to English Vowels
The RP accent of English has 20 vowel sounds, including monophthongs (short and long vowels) and diphthongs, but there is a discrepancy about the number of vowels in other varieties of English. Vowels can be transcribed in many different ways because accents differ greatly in the vowels they use, and there is no single right way of transcribing even one accent. The difference in English vowels is not only related to the number of vowels but also to the 'length' and 'quality' of vowel sounds. To fully understand English vowels, we must examine various varieties of English as well as vowel quality and vowel space.
Topic-071: Vowel Quality
Quality is a term used in auditory phonetics and phonology to refer to the characteristic resonance, or timbre of a sound, which results from the range of frequencies constituting the sound’s identity. Variations in vowels are describable in terms of quality; for example, the distinction between [i] and [e] is a qualitative difference. A major problem in describing vowels is the difficulty in precisely describing the tongue position during production, as people cannot easily determine where their tongues are. It is important to remember that the terms used for vowel description are simply labels that describe how vowels sound in relation to one another, not absolute descriptions of tongue position. This is because it is possible to make a vowel sound halfway between a high-vowel and a mid-vowel, or at any specified distance between any two other vowels, as vowels form a continuum. For example, gliding from /æ/ in had to /i/ in he demonstrates the difference in vowel quality.
🔑 Definition — Vowel Quality: The characteristic resonance or timbre of a vowel sound, determined by the range of frequencies that make up its identity, distinguishing it from other vowels (e.g., the difference between [i] and [e]). 💡 Why this matters: Recognizing that vowel quality is an auditory continuum, not a set of discrete tongue positions, is crucial for understanding how phoneticians describe and transcribe vowels across languages.
Topic-072: Auditory Vowel Space
Vowel sounds are tricky to describe phonetically accurately because they are points, or rather areas, within a continuous space called auditory vowel space. A language has a finite number of contrasting vowels, each represented by a discrete alphabetic symbol, but phonetically each corresponds to a range of typical values. Between any two actual vowel sounds, there is a gradient continuum that determines the dimensions of this space. Phonetically, the four vowels [i, æ, ɑ, u] (as given in the cardinal vowel system) give us something like the four corners of a space showing the auditory qualities of auditory vowel space. Phoneticians often use terms like high, low, back, and front when they simply label the auditory qualities of vowels and do not describe tongue positions.
📐 Definition: Auditory Vowel Space: A continuous, two-dimensional conceptual space representing the range of possible vowel qualities, with cardinal vowels like [i, æ, ɑ, u] acting as reference points at its corners. 💡 Why this matters: The concept of vowel space explains how languages can have many vowel sounds that are systematically ordered, and why descriptions like "high front vowel" are auditory labels, not precise tongue measurements.
Topic-073: American and British Vowels
Many American vowels are different from those in British English, making it a different English (e.g., compare Standard American Newscaster English with BBC English). When listening to American vowels [i, ɪ, ɛ, æ] as in words heed, hid, head, had spoken by a native speaker, these vowels sound as if they differ by a series of equal steps. Some Eastern American speakers make a distinct diphthong in heed, so their [i] is really a glide starting from almost the same vowel as that in hid. Similarly, back vowels also vary considerably; many Californians do not distinguish between the vowels in words father and author. The vowels [ʊ, u] as in good and food also vary, with a very unrounded vowel in good and a rounded but central vowel in food. In short, American English is distinct from British English, and students of phonetics should explore these differences.
🔑 Definition — Diphthong: A vowel sound that begins at one vowel quality and glides to another within the same syllable, as in the American pronunciation of heed starting near the vowel of hid. 💡 Why this matters: Understanding dialectal variation in vowel production is essential for accurate phonetic transcription and for analyzing how different English accents systematically differ in their vowel systems.
⭐ Key Takeaways
This lecture establishes that English vowels are fundamentally characterized by quality, which is an auditory property of resonance, not a simple measure of tongue position. Vowels exist on a continuum within an auditory vowel space, with cardinal vowels providing reference points like the four corners of that space. A critical point is that the terms "high," "low," "front," and "back" are auditory labels, not precise anatomical descriptions of tongue placement. Finally, students must remember that varieties of English differ significantly in their vowel systems, as seen in the distinct qualities of American and British vowels, including differences in length, rounding, and diphthongization.
🧠 Quick Revision Questions
- Why is it problematic to describe vowels solely in terms of tongue position?
- What are the four cardinal vowels that form the "corners" of the auditory vowel space?
- How do phoneticians use terms like "high, low, back, and front" when describing vowels, if not for tongue position?
- Describe one specific difference between American and British vowels for the words heed or father/author.
- What does it mean to say that vowels form a "continuum"?
📘 Lecture 14 — English Vowels-II
📖 Overview: This lecture continues the study of English vowels by focusing on complex vowel types: diphthongs, triphthongs, and rhotic vowels. It also explains how vowel quality changes in stressed, unstressed, and reduced syllables. Understanding these concepts is essential for accurate pronunciation and phonetic transcription.
🗂️ Topics Covered
This lecture covers diphthongs, their classification as gliding vowels, and their comparison with monophthongs and triphthongs. It then introduces rhotic vowels in American and other English varieties. Finally, it explains the three forms a vowel can take (stressed, unstressed, reduced) and the role of the schwa vowel in reduced syllables.
📝 Lecture Summary
Topic-074: Diphthongs
A diphthong is a single vowel consisting of the features of two vowels. Its most important feature is the glide from one vowel quality to another (so basically it is a glide). The BBC accent of English contains a large number (eight in total) of diphthongs including three ending at /ɪ/ (eɪ, aɪ, ɔɪ – as in words bay, buy and boy), two ending at /ʊ/ (əʊ, aʊ – as in words no and now) and three ending at /ə/ (ɪə, eə, ʊə - as in words peer, pair and poor). There had been a point of difference whether a diphthong should be treated as a single phoneme (in its own right) or it is a combination of two phonemes.
🔑 Definition — Diphthong: a vowel where there is a single (perceptual) noticeable change in quality during a syllable (as in English words beer, time and loud). 🔑 Definition — Monophthong: a vowel with no qualitative change in it. 🔑 Definition — Triphthong: a vowel where two such changes can be heard.
Diphthongs, or ‘gliding vowels’, are usually classified into phonetic types depending on one of the two elements that is the more sonorous. ‘Falling’ (or ‘descending’) diphthongs have the first element stressed. In the English examples: ‘rising’ (or ‘ascending’) diphthongs have the second element stressed.
💡 Why this matters: Knowing which element is stressed in a diphthong (falling vs. rising) helps predict the sound pattern and rhythm of words in English.
Topic-075: Rhotic Vowels
This term is used to describe some varieties of English (e.g., American) pronunciation in which the /r/ phoneme is found in all its phonological contexts. Remember that in the BBC accent of English, /r/ is only found before vowels (as in ‘red’ /red/, ‘around’ /əraʊnd/), but never before consonants or before a pause. In rhotic (e.g., some American) accents, on the other hand, /r/ may occur before consonants (as in ‘cart’ /ka:rt/) and before a pause (as in ‘car’ /kɑ:r/). While the BBC accent is non-rhotic, many accents of the British Isles are rhotic (including most of the south and west of England, much of Wales, and all of Scotland and Ireland). Similarly, most speakers of American English speak with a rhotic accent, but there are non-rhotic areas including the Boston area, lower-class New York and the Deep South. From English language teaching point of view, foreign learners encounter a lot of difficulty in learning not to pronounce /r/ in the wrong places.
🔑 Definition — Rhotic accent: a variety of English in which the /r/ phoneme is found in all its phonological contexts, including before consonants and before a pause. 🔑 Definition — Non-rhotic accent: a variety of English (like BBC) in which /r/ is only found before vowels.
📌 Example: In a rhotic accent, ‘car’ is pronounced /kɑ:r/ (with the /r/ sounded); in a non-rhotic accent, it is pronounced /kɑ:/ (without the /r/).
Topic-076: Unstressed Syllables
A vowel may take one out of three forms: stressed, unstressed and reduced. Most of the time a vowel is completely pronounced when it is in a stressed syllable but the same vowel is different in quality (allophonic form) when it takes place in an unstressed syllable, and, of course, it is reduced to another form when it is in a reduced syllable. Remember that in most cases, various reduced vowels are taking the shape of a schwa vowel /ə/. The symbol /ə/ may be used to show many types of vowels with a central, reduced vowel quality. A vowel in an unstressed syllable does not necessarily have a completely reduced quality. All the English vowels can occur in unstressed syllables in their full, unreduced forms and not all but many of them can occur in all possible three forms.
🔑 Definition — Reduced vowel: a vowel that changes to a central, reduced quality (usually schwa /ə/) when in an unstressed or reduced syllable. 📐 Formula: Stressed vowel → [full quality] / Unstressed vowel → [can be full or allophonic] / Reduced vowel → [schwa /ə/]
📌 Example: The vowel in 'about' /əˈbaʊt/ is a reduced schwa; in 'photograph' /ˈfəʊ.tə.ɡrɑːf/, the first vowel is stressed and full, while the second is unstressed and reduced to schwa.
⭐ Key Takeaways
This lecture is critical for understanding complex vowel sounds in English. A diphthong is a single vowel with a glide from one quality to another, and it is classified as falling (first element stressed) or rising (second element stressed). The BBC accent has eight diphthongs, but rhotic accents like American English pronounce /r/ in all positions, unlike non-rhotic BBC English. Vowels can appear in three forms: stressed (full quality), unstressed (may be full or allophonic), and reduced (typically becoming schwa /ə/). Foreign learners must be careful not to insert /r/ in the wrong places when learning non-rhotic accents.
🧠 Quick Revision Questions
- What is the defining characteristic of a diphthong compared to a monophthong?
- How many diphthongs does the BBC accent of English contain? List the three groups based on their ending sound.
- What is the difference between a rhotic and a non-rhotic accent?
- Give an example of a word where /r/ would be pronounced in a rhotic accent but not in a non-rhotic accent.
- What are the three forms a vowel can take, and which form is most commonly associated with the schwa /ə/?
📘 Lecture 15 — English Vowels-I II
📖 Overview: This lecture introduces the fundamental distinction between tense and lax vowels, along with their consonantal counterparts fortis and lenis, which describe articulatory strength. It also provides a comprehensive set of rules governing English vowel allophones, explaining how vowel length, quality, and voicing vary depending on phonetic context. Understanding these concepts is critical for accurate phonetic transcription and phonological analysis.
🗂️ Topics Covered
The lecture covers three main topics: first, the distinction between tense and lax vowels based on muscular effort and duration; second, the analogous classification of fortis and lenis consonants based on articulatory strength and breath force; and third, six specific rules for English vowel allophones that govern variations in vowel length, stress effects, syllable count effects, voicelessness, nasalization, and retraction before certain consonants.
📝 Lecture Summary
Topic-077: Tense and Lax Vowels
This section explains the distinction between tense and lax vowels, which are labels for ‘strong’ and ‘weak’ vowels based on their behavior. This is a comparative feature from Jakobson and Halle’s distinctive feature theory in phonology. Lax sounds are produced with less muscular effort and movement, are relatively short, and include vowels like /ɪ, e, ɒ, æ, ʌ, ʊ, ə/ — vowels articulated near the center of the vowel area. In contrast, tense sounds involve greater articulatory energy and include vowels like /uː, iː, ɜː, aː, ʊə, iə/. Since there is no established standard for measuring articulatory energy, these terms only have meaning in relation to each other. The terms are mainly used by American phonologists when describing English vowels, and they can also apply to consonants as equivalents to fortis (tense) and lenis (lax), though this is not common today.
🔑 Definition — Lax vowel: A vowel produced with relatively little articulatory energy, relatively short and indistinct, typically articulated near the center of the vowel area.
🔑 Definition — Tense vowel: A vowel produced with a relatively greater amount of articulatory energy, typically longer and more peripheral in articulation.
💡 Why this matters: The tense-lax distinction helps explain many phonological patterns in English, including vowel length differences and the behavior of vowels before certain consonants.
Topic-078: Fortis and Lenis Consonants
This section describes fortis and lenis as terms used in the phonetic classification of consonantal sounds based on their manners of articulation. Fortis refers to a sound made with a relatively strong degree of muscular effort and breath force, while lenis refers to its weaker counterpart. The distinction between tense and lax is used for vowels on similar lines. The labels ‘strong’ and ‘weak’ are sometimes used but are prone to ambiguity. In English, voiceless consonants like /p, t, f, s/ tend to be produced with fortis articulation, while their voiced counterparts are relatively weak or lenis. When the voicing distinction is reduced (e.g., in whispered speech or certain contexts), it is only the degree of articulatory strength that maintains the contrast between sounds. The term ‘fortis’ is occasionally used loosely for strong vowel articulation, but this is not standard practice.
🔑 Definition — Fortis: A consonant sound made with a relatively strong degree of muscular effort and breath force (e.g., /p, t, f, s/ in English).
🔑 Definition — Lenis: A consonant sound made with a relatively weak degree of muscular effort and breath force (e.g., /b, d, v, z/ in English).
📌 Example: In the pair “sip” /sɪp/ vs. “zip” /zɪp/, the /s/ is fortis (voiceless, strong) and the /z/ is lenis (voiced, weak). If voicing is removed, the strength difference still distinguishes them.
Topic-079: Rules for English Vowel Allophones
This section presents six specific rules that govern how English vowels change in different phonetic environments. These rules describe systematic variations in vowel length, quality, and voicing that are predictable based on context.
Rule 1: Other things being equal, a given vowel is longest in an open syllable, next longest in a syllable closed by a voiced consonant, and shortest in a syllable closed by a voiceless consonant.
📌 Example: Compare “sea” /siː/ (open syllable, longest vowel), “seed” /siːd/ (closed by voiced /d/, medium length), and “seat” /siːt/ (closed by voiceless /t/, shortest vowel). Similarly: “sigh” /saɪ/, “side” /saɪd/, “site” /saɪt/.
Rule 2: Other things being equal, vowels are longer in stressed syllables than in unstressed syllables.
📌 Example: Compare “below” /bɪˈləʊ/ (stressed second syllable has longer vowel) and “billow” /ˈbɪləʊ/ (stressed first syllable has longer vowel).
Rule 3: Other things being equal, vowels are longest in monosyllabic words, next longest in words with two syllables, and shortest in words with more than two syllables.
📌 Example: Compare “speed” /spiːd/ (one syllable, longest vowel), “speedy” /ˈspiːdi/ (two syllables, medium), and “speedily” /ˈspiːdɪli/ (three syllables, shortest vowel).
Rule 4: A reduced vowel may be voiceless when it is after a voiceless stop (and before a voiceless stop).
📌 Example: Compare “potato” /pəˈteɪtəʊ/ (the reduced vowel after /p/ may devoice) with “catastrophe” /kəˈtæstrəfi/ (similar devoicing pattern).
Rule 5: Vowels are nasalized in syllables closed by a nasal consonant.
📌 Example: In the word “man” /mæn/, the vowel /æ/ is nasalized because it precedes the nasal /n/. The tilde diacritic [˜] can be used to indicate this: [mæ̃n].
Rule 6: Vowels are retracted before syllable-final dark [l̴].
📌 Example: Compare the pronunciation of /iː/ in “heed” (non-retracted) vs. “heel” (retracted before [l̴]); /eɪ/ in “paid” vs. “pail”; and /æ/ in “pad” vs. “pal”. The vowel in “heel,” “pail,” and “pal” is pulled back in the mouth compared to the versions without final [l̴].
💡 Why this matters: These six rules explain why English vowels sound different in different words and contexts. They are essential for accurate phonetic transcription and for understanding why native speakers produce systematic variations in pronunciation.
⭐ Key Takeaways
The most critical points to remember are: (1) Tense vowels involve greater muscular effort and are longer than lax vowels, which are weaker and shorter; (2) Fortis consonants are strong/voiceless and lenis consonants are weak/voiced, and this strength distinction can maintain contrast even when voicing is neutralized; (3) Vowel length varies systematically — longest in open syllables, medium before voiced consonants, and shortest before voiceless consonants; (4) Vowels are longer in stressed syllables and in shorter words; and (5) Vowels become nasalized before nasal consonants and retracted before dark [l̴], while reduced vowels can become voiceless after voiceless stops. These allophonic rules are predictable and must be applied when transcribing spoken English phonetically.
🧠 Quick Revision Questions
- What is the primary difference between a tense vowel and a lax vowel in terms of articulation and duration?
- How do the terms fortis and lenis relate to the tense/lax distinction, and which English consonants typically fall into each category?
- According to Rule 1 for vowel allophones, how does vowel length compare in the words “bee,” “bead,” and “beat”? Explain why.
- In what phonetic environment do vowels become nasalized in English, and what process causes this change?
- Predict the allophonic variations in the word “spoon” /spuːn/ — considering stress, syllable structure, and the final consonant.
📘 Lecture 16 — English Words and Sentences-I
📖 Overview: This lecture examines how words change their pronunciation when moving from isolation to connected speech. It introduces the critical distinction between strong and weak forms of grammatical words, defines stress and its degrees, and explains how stress operates at both word and sentence levels to create meaning and rhythm in English.
🗂️ Topics Covered
The lecture begins by contrasting citation speech versus connected speech, focusing on how closed-class words (grammatical words) typically appear in weak forms. It then explains the two possible pronunciations for words (strong and weak forms), defines stress as prominence, and discusses different degrees of stress (primary, secondary, tertiary, weak). Finally, it covers how stress functions at lexical and sentence levels to change grammatical categories and create rhythm.
📝 Lecture Summary
Topic-080: English Words and Sentences
There is a lot of difference between words spoken in isolation versus in connected speech. The key difference between citation speech (where a word is in its complete form) and connected speech is the variable degree of emphasis placed on different words. This “degree of emphasis” relates to the amount of information a word conveys. The difference is particularly noticeable for the closed class of words — grammatical words such as determiners (a, an, the), conjunctions (and, or), and prepositions (of, in, with). These are very rarely emphasized in connected speech, so their normal pronunciation differs from their citation forms. However, closed-class words show a strong form when emphasized (e.g., He wanted pie and ice cream, not pie or ice cream) and a weak form when unstressed.
🔑 Definition — Citation speech: The complete, full form of a word spoken in isolation, without reductions. 🔑 Definition — Connected speech: Natural, flowing speech where words undergo pronunciation changes due to context and stress patterns. 🔑 Definition — Closed class words: Grammatical words (determiners, conjunctions, prepositions) that rarely receive emphasis in connected speech.
Topic-081: Words in Connected Speech
Words can have two possible forms: weak and strong. The strong form occurs when a word is stressed (e.g., I want bacon and eggs where "and" is emphasized). The notion is also used for syntactically conditioned alternatives (e.g., your book vs. the book is yours). The weak form results when a word is unstressed, as in the normal pronunciation of "of" in cup of tea. Several closed class/function words in English have more than one weak form — for example, and [ænd] can be [ənd], [ən], [n], etc.
📌 Example: The word "and" has multiple weak forms: strong form [ænd], weak forms [ənd], [ən], [n].
Topic-082: Stress
Stress is a term used in phonetics to refer to the degree of force (making a syllable louder and longer) used in producing a syllable. The usual distinction is between stressed and unstressed syllables, the former being more prominent (marked in transcription with a raised vertical line [ˈ]). Prominence is due to an increase in loudness, but increases in length and often pitch also contribute. Stressed syllables are produced with greater effort and tend to be longer than unstressed. In terms of linguistic function, stress is treated under two headings: word stress (lexical stress) and sentence stress (emphatic stress).
🔑 Definition — Stress: The degree of force used in producing a syllable, making it more prominent through loudness, length, and pitch. 💡 Why this matters: Stress helps listeners identify key information and distinguishes words (e.g., record as noun vs. verb).
Topic-083: Degree of Stress
The analysis of the degree of stress attracted attention in the mid-twentieth century. The question is how many degrees of stress need to be recognized. In the American structuralist tradition, four degrees are usually distinguished as stress phonemes: (1) primary, (2) secondary, (3) tertiary, and (4) weak (from strongest to weakest). These contrasts are demonstrable only on words in isolation (e.g., elevator operator). In most phonological analysis, experts distinguish among three degrees: primary, secondary, and weak (or unstressed) — e.g., /ɪg.ˌzæm.ɪ.ˈneɪ.ʃən/.
📐 Formula: Four degrees (structuralist): Primary > Secondary > Tertiary > Weak 📐 Formula: Three degrees (phonological): Primary > Secondary > Weak/Unstressed 📌 Example: The word examination /ɪg.ˌzæm.ɪ.ˈneɪ.ʃən/ shows three degrees — primary stress on [-neɪ-], secondary on [-zæm-], weak on other syllables.
Topic-084: Stress Explanation
Stress is a large topic with many areas of disagreement. Stress is basically a prominence of syllable in terms of loudness, length, pitch, and quality, all working together. Two types of stress are important: lexical stress (stress on a syllable within a word), which changes grammatical category (compare inˈsult (verb) with ˈinsult (noun)) and meaning; and sentence level or prosodic stress (stress on certain words within a sentence), which is a change of ‘beat’ on certain words. We create rhythm in spoken language based on stress. The distinguishing degree of emphasis is used for creating contrast and is part of language formality and intonation.
📌 Example: Mary’s younger brother wanted fifty chocolate peanuts. (stressed words in bold show the natural rhythm pattern)
⭐ Key Takeaways
The critical distinction between citation speech and connected speech explains why grammatical words (closed class) typically appear in weak forms when unstressed, with multiple possible weak forms (e.g., and → [ən], [n]). Stress is defined as prominence through loudness, length, pitch, and quality, and can be analyzed as either four degrees (American structuralist) or three degrees (most phonological analysis). Lexical stress changes word category and meaning (e.g., record), while sentence stress creates rhythm and emphasis in connected speech. Understanding weak and strong forms is essential for natural pronunciation and listening comprehension in English.
🧠 Quick Revision Questions
- What is the difference between citation speech and connected speech, and which class of words shows the most noticeable difference?
- How many weak forms can the word "and" have, and what are some examples?
- List the four degrees of stress according to the American structuralist tradition, from strongest to weakest.
- What phonetic features contribute to making a syllable stressed?
- How does lexical stress differ from sentence stress, and provide one example of each.
📘 Lecture 17 — English Words and Sentences-II
📖 Overview: This lecture examines the rhythmic and melodic features of connected English speech. It explains sentence rhythm as a perceptual phenomenon, introduces intonation as pitch variation that conveys meaning and attitude, and defines target tones as linguistically contrastive pitch movements. Understanding these suprasegmental features is essential for analyzing how English speakers organize and interpret spoken language beyond individual sounds.
🗂️ Topics Covered
The lecture covers sentence rhythm and the stress-timed versus syllable-timed language distinction, intonation as pitch variation at the sentence level, the analysis of intonation through tonic accents and tone units, the linguistic and attitudinal functions of intonation, the definition and analysis of target tones in tonal and non-tonal languages, and the use of the ToBI (Tone and Break Indices) system for describing intonational patterns.
📝 Lecture Summary
Topic-085: Sentence Rhythm
Sentence rhythm refers to the way speech events are distributed in time. While obvious examples exist in chanting (e.g., children skipping or cricket crowds), conversational speech rhythm is more complex but not random. The stress-timed rhythm hypothesis proposes that English speech can be divided into roughly equal time intervals called feet, each beginning with a stressed syllable. In contrast, syllable-timed languages have syllables of approximately equal duration regardless of stress. Evidence from real speech suggests such rhythms are mainly found in careful, controlled speech, but psychological research indicates that listeners’ brains tend to perceive timing regularities even when little physical regularity exists.
💡 Why this matters: The stress-timed vs. syllable-timed distinction helps explain why English speakers compress unstressed syllables, making them harder for learners to hear, while speakers of syllable-timed languages (like Spanish or French) give each syllable more equal duration.
Topic-086: Intonation
Intonation refers to variations in the pitch of a speaker’s voice (f₀) used to convey or alter meaning. In its broader sense, intonation covers much of the field of prosody, including variations in voice quality, tempo, and loudness. Pitch movement can be analyzed to find regular patterns. Some experts look for an underlying basic pitch melody (or a small number of melodies) and describe deviations from these. Others break pitch patterns into small constituent units such as pitch phonemes and pitch morphemes. The most widely used British approach takes the tone unit as its basic unit and examines pitch possibilities of its components: pre-head, head, tonic syllable/nucleus, and tail. Intonation conveys emotions and attitudes and serves other linguistic functions, including signaling grammatical structure and new information through prominence. Interesting relationships exist between intonation and grammar—for example, a perceived difference in grammatical meaning may depend on pitch movement.
Topic-087: Explaining Intonation
Intonation is pitch variation at sentence level and can be described in terms of the intonational phrase. To describe intonation, we analyze the role of a stressed syllable whose pitch change creates a major change called the tonic accent (marked with an asterisk), forming the pitch peak in an intonational phrase. A formal category of intonational phrase is sometimes recognized as an utterance span dominated by boundary tones. Intonation performs several functions:
- Grammatical function: It signals grammatical structure, similar to punctuation in writing (marking sentence, clause, and other boundaries). It contrasts grammatical structures like questions and statements. For example, ‘He’s going, isn’t he?’ (rising pitch = asking) vs. ‘He’s going, isn’t he!’ (falling pitch = telling).
- Attitudinal function: It communicates personal attitude (e.g., sarcasm, puzzlement, anger) through contrasts in pitch along with other prosodic and paralinguistic features.
- Social function: It may signal social background.
🔑 Definition — Tonic accent: The syllable in an intonational phrase that carries the major pitch movement or pitch peak, marked with an asterisk.
📌 Example: ‘He’s going, isn’t he?’ with rising pitch on ‘he’ signals a genuine question, while ‘He’s going, isn’t he!’ with falling pitch signals a statement expecting agreement.
Topic-088: Target Tones
In phonetics and phonology, tone has a restricted meaning: it refers to an identifiable movement or level of pitch used in a linguistically contrastive way. In tone languages, tone changes the meaning of a word. For example, in Mandarin Chinese, /má/ said with high pitch means ‘mother’, while /mǎ/ spoken on a low rising tone means ‘hemp’. In non-tonal languages like English, tone forms the central part of intonation—the difference between a rising and falling tone on a particular word may cause a different interpretation of the sentence. In tone languages, tones are properties of individual syllables, whereas an intonational tone may be spread over many syllables. In English intonation analysis, tone refers to one of the pitch possibilities for the tonic (or nuclear) syllable, a set usually including fall, rise, fall–rise, and rise–fall, though others are suggested by various experts.
📌 Example (Mandarin): /má/ (high pitch) = ‘mother’; /mǎ/ (low rising pitch) = ‘hemp’
📌 Example (English): A rising tone on “really” in “You’re coming, really?” suggests surprise or doubt, while a falling tone suggests acceptance.
Topic-089: Explaining Target Tones
Several approaches analyze intonation. Some describe pitch patterns as contours analyzed in terms of pitch levels as pitch phonemes and morphemes. Others describe patterns as tone units or tone groups, analyzed as contrasts of nuclear tone and tonicity. Three variables are generally distinguished: (1) pitch range, (2) height, and (3) direction. Some approaches, especially within pragmatics, operate with a broader notion than the tone unit, viewing intonational phrasing as a structured hierarchy of intonational constituents in conversation. A formal category of intonational phrase is also sometimes recognized as an utterance span dominated by boundary tones. One recently developed method is ToBI (Tone and Break Indices), used for describing intonation by representing High (H) and Low (L) pitches in a sentence, showing pitch accent, phrase accent, and boundary (of the phrase) through tone and break indices.
🔑 Definition — ToBI (Tone and Break Indices): A system for describing intonation using High (H) and Low (L) pitch targets to represent pitch accents, phrase accents, and boundary tones, along with break indices indicating degrees of juncture between words.
⭐ Key Takeaways
Sentence rhythm in English is described by the stress-timed rhythm hypothesis, where speech is perceived in roughly equal intervals called feet beginning with stressed syllables—this contrasts with syllable-timed languages. Intonation is pitch variation at sentence level that serves grammatical functions (like distinguishing questions from statements), attitudinal functions (conveying emotion), and social functions. Target tones in English are one of a set of pitch possibilities for the tonic syllable (fall, rise, fall–rise, rise–fall), whereas in tone languages like Mandarin, tones change word meanings. The ToBI system provides a standardized method for describing intonation using High/Low pitch targets and break indices. Understanding these suprasegmental features is essential for analyzing natural speech beyond individual sounds.
🧠 Quick Revision Questions
- What is the stress-timed rhythm hypothesis, and how does it differ from syllable-timed languages?
- List and briefly explain the components of a tone unit in the British approach to intonation analysis.
- How does intonation distinguish between ‘He’s going, isn’t he?’ as a question versus a statement?
- What is the difference between tone in a tone language (e.g., Mandarin) and tone in English intonation?
- What does the ToBI system use to represent intonational patterns, and what are its main components?
📘 Lecture 18 — AIRSTREAM MECHANISMS
📖 Overview: This lecture explains the physiological processes that provide the energy source for all human speech sounds. It details the three major airstream mechanisms—pulmonic, glottalic, and velaric—and describes how each mechanism initiates airflow (either egressive or ingressive) to produce specific types of consonants found across the world's languages.
🗂️ Topics Covered
The lecture begins by defining airstream mechanisms and the two directions of airflow (egressive and ingressive). It then systematically explains the pulmonic airstream mechanism, the most common for speech, using the lungs and diaphragm. Next, it covers the glottalic mechanism, produced by moving the larynx with closed vocal folds, giving rise to ejective and implosive sounds. Finally, it describes the velaric mechanism, which creates clicks by sucking air using the back of the tongue against the velum, and concludes with a summary comparing all three mechanisms and their linguistic occurrences.
📝 Lecture Summary
Topic-090: Airstream Mechanisms
All human speech sounds are produced by making the air move in the oral and nasal cavity, creating an airstream. The study of how and what type of air moves is called the airstream mechanism. Most commonly, air is moved outwards from the body, creating an egressive airstream; more rarely, sounds are made by drawing air inward, creating an ingressive airstream. The airstream provides the source of energy for speech sound production. There are three main mechanisms: the pulmonic airstream (using the lungs), the glottalic airstream (using the movement of the glottis), and the velaric airstream (using the back of the tongue against the velum).
🔑 Definition — Airstream Mechanism: the physiological process that initiates the movement of air in the vocal tract, providing the energy for speech sound production.
Topic-091: Pulmonic Airstream Mechanism
The pulmonic airstream mechanism is the most commonly used mechanism for speech production. Almost all sounds we produce in speaking are created with the help of air compressed by the lungs. The adjective 'pulmonic' refers to this lung-created airstream. For speaking, the pulmonic airstream is always egressive (speech sounds are produced while pushing the air out), although it may be ingressive when breathing in. The mechanism involves the human respiratory system: the respiratory muscles set the air in motion, and the lungs (sponge-like tissues) are contained within an air cage called the diaphragm. The diaphragm contracts and enlarges the lung cavity, creating egressive and ingressive actions. This mechanism sets an airflow for speech production, and human beings produce speech sounds while pushing the air out.
🔑 Definition — Pulmonic Airstream: an airstream initiated by the lungs and respiratory muscles; in speech, it is almost always egressive.
Topic-092: Glottalic Airstream Mechanism
This mechanism involves the glottis (the aperture between the vocal folds). A glottalic airstream is produced by making a tight closure of the vocal folds and then moving the larynx up or down. Raising the larynx pushes the air outwards, causing an egressive glottalic airstream; lowering the larynx pulls air into the vocal tract, causing an ingressive glottalic airstream. Sounds produced this way are called ejective (egressive) or implosive (ingressive), respectively. Glottalization is a process used for any articulation involving a simultaneous glottal constriction (e.g., a glottal stop). In English, glottal stops often reinforce a voiceless plosive at the end of a word, as in what. These sounds are made while the glottis is closed, without direct involvement of air from the lungs. Air is compressed in the mouth or pharynx above the glottal closure and released while the breath is held; the resultant sounds are called ejective sounds (also called glottalic or glottalized sounds, though the latter term often refers to secondary articulation). In languages like Quechua and Hausa, ejective consonants are used as phonemes. A further category of sounds involving this mechanism is known as implosive (ingressive glottalic).
💡 Why this matters: Glottalic sounds are a common feature in many non-European languages and understanding them is crucial for cross-linguistic phonetics.
🔑 Definition — Ejective: a sound produced by an egressive glottalic airstream, created by raising the closed glottis. 🔑 Definition — Implosive: a sound produced by an ingressive glottalic airstream, created by lowering the closed glottis.
Topic-093: Velaric Airstream Mechanism
The velaric airstream mechanism involves the velum (soft palate). Under this mechanism, speech sounds are made by sucking the air inward. This sucking mechanism is used first by babies for feeding and later by adults for actions like sucking liquid through a straw or drawing smoke from a cigarette, using the back of the tongue against the velum. The basic mechanism requires an air-tight closure between the back of the tongue and the soft palate. The tongue is then retracted, lowering pressure in the oral cavity, and suction takes place. Consonants produced with this mechanism are called clicks. These sounds have a distinctive role in some languages, such as Zulu. In English, they may be heard in the 'tut tut' (or tsk tsk) sounds of disapproval and in a few other contexts.
🔑 Definition — Velaric Airstream: an airstream initiated by sucking action, using a closure between the back of the tongue and the velum. 🔑 Definition — Click: a consonant produced with a velaric airstream mechanism, involving an ingressive airflow.
Topic-094: Summary of the Airstream Mechanisms
There are three possible mechanisms for human speech production. The most common is the pulmonic airstream, usually an egressive one produced by compressing the lungs and expelling air through the vocal tract (occasionally speech is produced while breathing in). The second is the glottalic mechanism, produced by the larynx with closed vocal folds, moved up and down like the plunger of a bicycle pump. The last is the velaric mechanism, where the back of the tongue is pressed against the soft palate, making an air-tight seal, and then drawn backwards or forwards to produce an airstream. The ingressive glottalic consonants (implosives) and egressive ones (ejectives) are found in many non-European languages. Click sounds (ingressive velaric) are rarer but occur in southern African languages like Nàmá, Xhosa (or Hausa), and Zulu. Speakers of other languages, including English, use click sounds for non-linguistic communication, as in the 'tut-tut' (tsk-tsk) sound of disapproval.
💡 Why this matters: This summary provides a clear comparative framework for distinguishing the three airstream mechanisms and their linguistic distributions.
⭐ Key Takeaways
The most critical concept from this lecture is that all speech sounds require an airstream as their source of energy, and there are exactly three types of airstream mechanisms: pulmonic, glottalic, and velaric. The pulmonic mechanism, using the lungs, is the most common and is almost always egressive for speech. The glottalic mechanism involves moving the larynx with closed vocal folds, producing ejective (egressive) and implosive (ingressive) sounds. The velaric mechanism uses a tongue-velum closure to create suction, producing click sounds. For the exam, remember the specific articulators involved, the direction of airflow (egressive vs. ingressive) for each mechanism, and the names of the resulting sound types (ejectives, implosives, clicks).
🧠 Quick Revision Questions
- What are the two possible directions of airflow in an airstream mechanism, and which one is most common in speech?
- Which airstream mechanism is used for the majority of speech sounds in the world's languages, and what is its primary initiator?
- How is an ejective sound produced in the glottalic airstream mechanism, and what is the movement of the larynx?
- What is the defining characteristic of a click sound's production in the velaric airstream mechanism?
- Name one language that uses ejective consonants as phonemes and one language that uses clicks as phonemes.
📘 Lecture 19 — PHONATION
📖 Overview: This lecture explores the process of phonation, which describes the forms of vibration of the vocal folds (voicing) within the larynx. It explains how different states of the glottis create distinct voice qualities such as voiced, voiceless, creaky, and breathy sounds, and emphasizes how voicing distinguishes consonants and changes word meanings.
🗂️ Topics Covered
This lecture introduces phonation as the technical term for voicing and laryngeal activity, explaining the role of the larynx and vocal folds. It details the states of the glottis that produce different voicing features, describes four main phonation types (voiceless, voiced, creaky, and breathy), and examines how voicing functions as a distinctive feature for consonants, particularly obstruents.
📝 Lecture Summary
Topic-095: Introduction to Phonation
The position of the larynx (sound box) and the vocal folds inside it are crucial for describing speech sounds. Phonation is the technical term for the forms of vibration of the vocal folds, more commonly known as voicing. The glottis (the space between the vocal folds) can assume various shapes, such as voiced, voiceless, murmuring, and creaky positions. The most common positions describe consonants as either voiceless (vocal folds apart, e.g., /p/, /t/) or voiced (folds nearly together and vibrating, e.g., /b/, /g/). These glottal states are important for describing sounds in specific languages and pathological voices.
🔑 Definition — Phonation: The technical term for describing the forms of vibration of the vocal folds (or vocal cords); the process is more commonly known as voicing. 🔑 Definition — Glottis: The space between the vocal folds.
Topic-096: Phonation Explanation
Phonation is a general term in phonetics for any vocal activity in the larynx. The main phonatory activities are the various kinds of vocal-fold vibration (voicing or phonation). The study of phonation types accounts for laryngeal possibilities like voiced, voiceless, breathy, and creaky voice. Some phoneticians include modifications from variations in length, thickness, and tension of the vocal folds, seen in different speech registers. During phonation, air passes between the vocal folds, and modification to the air passage creates variation in intensity, frequency, and quality of the sound, playing an important role in voicing and murmuring.
💡 Why this matters: Understanding phonation explains how the larynx physically modifies air to produce different voice qualities and sound variations.
Topic-097: States of the Glottis
The space between the vocal folds (glottis) can assume a number of positions, modifying speech sounds. The contact of vocal folds and the shape of vibration produce phonation or voicing. Based on the nature of vibration, the states of the glottis are determined, causing differences in pitch, loudness, and voice quality. A narrow opening between the vocal folds produces friction noise, found in whispering and the glottal fricative /h/. A more widely open glottis is found in most voiceless consonants. Four possible states of the glottis are listed in Table 6.6 of the course book.
Topic-098: Phonation Types
There are four main possible glottis/larynx settings or types of phonation:
- Voiceless – Folds are open apart, air passes freely through the glottis (e.g., /t/, /p/).
- Voiced – Folds are tight together and vibrate during air passage through the glottis (e.g., /b/, /d/).
- Creaky voice – Slight opening in the front, arytenoid cartilages are tight together, so vocal folds vibrate only at the anterior end (small opening at the top).
- Breathy or murmuring sound – Vocal folds are apart but still vibrating; a breathy voice is like a whisper except with voice.
🔑 Definition — Creaky voice: A phonation type where there is a slight opening in the front and the arytenoid cartilages are tight together, allowing vocal folds to vibrate only at the anterior end. 🔑 Definition — Breathy voice: A phonation type where vocal folds are apart but still vibrating, producing a sound like a whisper except with voice.
Topic-099: Voicing and Consonants
Voicing is an important feature of speech sounds used for description and distinction. Sonorants (vowels, nasals, and approximants) are usually voiced, though voicing may be weak or absent in particular contexts. Obstruents (fricatives and plosives) may be voiced or voiceless and are the most frequently found sounds with both voicing and voicelessness. Voiced and voiceless distinctions can make a phonemic distinction in languages like English, changing the meaning of a word. A glottal stop is only an allophonic variation, used in RP before /p, t, k/, and in some dialects these can be replaced by a glottal stop. The symbol for a glottal stop is [ʔ].
🔑 Definition — Obstruents: A class of speech sounds (fricatives and plosives) that can be voiced or voiceless. 🔑 Definition — Glottal stop: An allophonic variation in RP occurring before /p, t, k/, represented by the symbol [ʔ].
⭐ Key Takeaways
Phonation is the technical term for voicing, describing the vibration of vocal folds in the larynx. The glottis can assume different states, producing four main phonation types: voiceless, voiced, creaky, and breathy. These states affect pitch, loudness, and voice quality, and are crucial for describing consonants. Voicing is a phonemic distinction in English, especially for obstruents (fricatives and plosives), where it can change word meaning. The glottal stop [ʔ] is an allophonic variation in RP, often before /p, t, k/.
🧠 Quick Revision Questions
- What is the technical term for the process of vocal fold vibration, and what is its more common name?
- Name the four main states of the glottis or types of phonation described in the lecture.
- Which class of speech sounds (obstruents or sonorants) most frequently shows both voiced and voiceless forms?
- What is the symbol for a glottal stop, and in what context does it appear in RP English?
- How does a voiceless sound differ from a voiced sound in terms of the position and activity of the vocal folds?
📘 Lecture 20 — Voice Onset Time (VOT)
📖 Overview: This lecture introduces Voice Onset Time (VOT), a key acoustic measure in phonetics that captures the timing of vocal-fold vibration relative to the release of a plosive (stop) consonant. VOT is crucial for distinguishing voicing and aspiration patterns across languages, and it provides a scientific basis for comparing stop sounds in different linguistic systems.
🗂️ Topics Covered
The lecture begins by defining Voice Onset Time (VOT) and explaining how it relates to voiced, voiceless, and voiceless aspirated plosives. It then discusses the importance of VOT in experimental phonology, language comparison, and bilingual perception, including examples from English and Navajo. Finally, it describes the three types of VOT: zero (unaspirated voiceless), positive (aspirated), and negative (voiced), with a figure illustrating the VOT continuum.
📝 Lecture Summary
Topic-100: Voice Onset Time (VOT)
All human languages distinguish between voiced and voiceless consonants, with plosives (stops) being the most common consonants to use this distinction. However, the timing of voicing relative to articulation is critical. In the case of aspiration, the beginning of full voicing is delayed after the release of a (usually voiceless) plosive. This delay (or lag) is measured scientifically as voice onset time (VOT). The onset of voicing may also precede the release, called a lead, resulting in fully or partially voiced plosives. Both cases are represented on the VOT scale: positive values (lag) and negative values (lead), with a third possibility of zero VOT.
🔑 Definition — Voice Onset Time (VOT): The point in time at which vocal-fold vibration starts in relation to the release of a closure during the production of plosive sounds. 📐 Formula: VOT = time of voicing onset − time of plosive release → Positive values = voicing after release (lag); negative values = voicing before release (lead); zero = voicing at release. 📌 Example: In English, the /p/ in “pin” has a positive VOT (approximately 50–60 ms), while the /b/ in “bin” may have a negative or zero VOT.
Topic-101: VOT Explanation
To understand VOT, three types of plosive sounds must be explained: voiced, voiceless unaspirated, and voiceless aspirated. During the production of a fully voiced plosive (e.g., /b/ or /g/), the vocal folds vibrate throughout. In a voiceless unaspirated plosive (e.g., /p/ or /t/), there is a delay (lag) before voicing starts. In a voiceless aspirated plosive (e.g., /pʰ/ or /tʰ/), the delay is much longer, depending on the amount of aspiration. This delay is called VOT, and it varies from language to language.
Topic-102: Importance of VOT
VOT is an important feature in experimental phonology used to analyze the nature of different languages and their stop sounds. Languages vary in VOT, and the delay (lag or VOT) is key for comparing languages. It also provides insight into perception of VOT by bilingual learners. Language-specific VOT values are important for language experts, and VOT values can reveal phonemic contrast in sound production by learners. VOT is calculated using a specific methodology, and contrastive features must be considered.
💡 Why this matters: VOT is not just a theoretical measure—it directly affects how speakers of different languages perceive and produce stop consonants, which is critical for language learning and cross-linguistic studies.
Different languages choose different points along the VOT continuum when forming oppositions among stop consonants. These possibilities are shown on a scale from most aspirated (largest positive VOT) to most voiced (largest negative VOT). For example, the Navajo aspirated stops have a very large VOT value, exceptional at 150 ms, while the normal VOT for English stressed initial /p/ is between 50 and 60 ms.
Topic-103. Types of VOT
There are three possible types of VOT based on the nature of stop sounds:
- Simple unaspirated voiceless stops: VOT is at or near zero — voicing of the following vowel begins at or near the release of the stop.
- Aspirated stops followed by a vowel: VOT is greater than zero, called positive VOT. The length of VOT corresponds to aspiration: longer VOT = stronger aspiration (e.g., Navajo at 150 ms, compared to English at 50–60 ms).
- Voiced stops: VOT is noticeably less than zero, called negative VOT — the vocal cords begin vibrating before the stop is released.
⭐ Key Takeaways
Voice Onset Time (VOT) is a fundamental acoustic measure that captures the timing of vocal-fold vibration relative to plosive release, allowing for a precise distinction among voiced, voiceless unaspirated, and voiceless aspirated stops. VOT values can be positive (lag), negative (lead), or near zero, and these values vary systematically across languages (e.g., Navajo aspirated stops have a large positive VOT of 150 ms, while English /p/ has about 50–60 ms). VOT is crucial for experimental phonology, language comparison, and bilingual speech perception, as it reveals phonemic contrasts and helps explain how different language systems organize stop consonant oppositions. The three VOT types—zero (unaspirated), positive (aspirated), and negative (voiced)—cover the full range of possibilities found in human languages.
🧠 Quick Revision Questions
- What does Voice Onset Time (VOT) measure, and how is it calculated?
- What are the three types of VOT, and what does each indicate about the nature of a stop sound?
- How does the VOT of English stressed initial /p/ compare to that of Navajo aspirated stops?
- Why is VOT important for studying bilingual learners?
- What does a negative VOT value indicate about the production of a plosive?
📘 Lecture 21 — Consonantal Gestures - I
📖 Overview: This lecture introduces the concept of consonantal gestures, treating speech sounds not as static points of articulation but as dynamic articulatory movements with inherent timing. It explains how experts describe and categorize unfamiliar sounds from various languages by focusing on articulatory targets and different types of gestures, including stops, nasals, and fricatives.
🗂️ Topics Covered
The lecture covers the definition of consonantal gestures and how sounds are treated as abstract movement patterns rather than exact locations. It then explores articulatory targets, explaining that traditional place names denote targets rather than fixed points. Several types of articulatory gestures are detailed, including bilabial, labiodental, dental, alveolar, retroflex, palato-alveolar/palatal, velar, pharyngeal, and epiglottal gestures. Finally, it provides a detailed examination of stops, nasals, and fricatives as gestural categories, with examples from English and other languages.
📝 Lecture Summary
Topic-104: Consonantal Gestures
In phonetics and phonology, speech sounds (segments) using basic units of contrast are defined as gestures – they are treated as the abstract characterizations of articulatory events with an intrinsic time dimension. Thus sounds are the underlying units represented by classes of functionally equivalent movement patterns. The purpose is to enhance the ability of experts to describe unfamiliar sounds of various languages, clarifying variation between varieties of a language and among different languages and language families. So involving the place and manners of articulation, sounds are treated as gestures rather than exact locations and manners.
Topic-105: Articulatory Targets
A number of possible places of articulation used in world languages have been defined so far. The traditional terms for places of articulation are not just names for particular locations; they should be thought of as names for articulatory targets. Many non-English sounds involve using gestures where the target or place of articulation is different from any similar sound found in English. For others, it is the type of gesture or the manner of articulation that is different. Experts illustrate these different types of targets by considering how each place of articulation is used in English and other languages for making stops, nasals, and fricatives.
Topic-106: Types of Articulatory Gestures
Based on the definition of gesture as a possibility of articulation of various speech sounds in contrast with English ones, the possible types are given here:
-
Bilabial gesture – very common in English (e.g., stops and nasal: p, b, m). In some languages (such as Ewe of West Africa), bilabial fricatives contrast with labiodental fricatives. The symbols for the voiceless and voiced bilabial fricatives are [ɸ, β]. These sounds are pronounced by bringing the two lips nearly together so there is only a slit between them.
-
Labiodental fricatives – many languages including English have the labiodental fricatives [f, v]. Probably no language has labiodental stops or nasals except as allophones of the corresponding bilabial sounds. In English, a labiodental nasal, [ɱ], may occur when /m/ occurs before /f/, as in emphasis or symphony.
-
Dental sounds are present in British and American English, e.g. dental fricatives [θ, ð], but there are no dental stops, nasals, or laterals except allophonically realized before [θ, ð] as in eighth, tenth, wealth. Many speakers of French, Italian, and languages like Urdu, Pashto, and Sindhi typically have dental stops such as [t̪, d̪].
-
Alveolar are very common targets – stops, nasals, and fricatives all occur in English and many other languages at alveolar as a target of articulatory gestures (e.g., t, d, n, l, r).
-
Retroflex is very common in many Pakistani languages, made by curling the tip of the tongue up and back so the tongue tip moves during sounds such as [ɳ, ŋ, ɲ, ʈ, ɽ]. These are also a feature of Indian English.
-
Palato-alveolar and palatal are also possible articulatory gestures. Similarly, velar sounds found in Urdu and other Pakistani languages include [x, ɣ] which are velar fricatives. The gestures for pharyngeal (such as Arabic pharyngeal fricative [ʕ]) and epiglottal sounds (such as epiglottal fricative [ʢ]) involve pulling the root of the tongue or the epiglottis back toward the back wall of the pharynx. These sounds are found in Arabic and other Semitic languages.
💡 Why this matters: Understanding these diverse gestural targets is crucial for accurately describing and transcribing sounds from languages beyond English.
Topic-107: Stops
Typical stop sounds are found in English and other languages, but there are interesting types found elsewhere. The following table shows how rich stop sounds are in various languages:
| Description | Symbol | Language |
|---|---|---|
| Voiced | b | English and other languages |
| Voiceless unaspirated | p | -do- |
| Aspirated | pʰ | Sindhi and many other Pakistani languages |
| Murmured (breathy) | bʱ | Sindhi |
| Implosive | ɓ | Sindhi |
| Laryngealized (creaky) | b̰ | Hausa |
| Ejective | kʼ | Hausa |
| Nasal release | dn | Russian |
| Prenasalized | nd | Swahili |
| Lateral release | tɬ | Navajo |
| Ejective lateral release | tɬʼ | Navajo |
| Affricate | ts | German |
| Ejective affricate | tsʼ | Navajo |
🔑 Definition — Stop: A speech sound produced by completely blocking the airflow in the vocal tract, which can vary in voicing, aspiration, and release type across languages.
Topic-108: Nasals
Nasal manners of articulation are commonly found in world languages. Like stops, nasals can occur voiced or voiceless (for example, in Burmese, Ukrainian, and French), though in English and most other languages nasals are voiced. Voiceless nasals are comparatively rare. They are symbolized by adding the voiceless diacritic [ ̥ ] under the symbol for the voiced sound. There are no special symbols for voiceless nasals; it is written as /m̥/ — a combination of the letter for the voiced bilabial nasal and a diacritic indicating voicelessness.
Topic-109: Fricatives
Fricative as an articulatory gesture may be divided into voiced or voiceless sounds, but we can also subdivide fricatives according to other aspects of the gestures that produce them. Some authorities divide fricatives into sounds like [s], where the tongue is grooved so the airstream comes through a narrow channel, and those like [θ], where the tongue is flat and forms a wide slit. A slightly better way is to separate them on a purely auditory basis. Say the English voiceless fricatives [f, θ, s, ʃ]. The two with the loudest high pitches are [s, ʃ]; they differ from [f, θ] in this way. The same difference occurs between [z, ʒ] and [v, ð]. The fricatives [s, z, ʒ, ʃ] are called sibilant sounds. They have more acoustic energy — greater loudness — at a higher pitch than other fricative sounds.
🔑 Definition — Sibilant: A type of fricative sound (e.g., [s, z, ʃ, ʒ]) characterized by greater acoustic energy and higher pitch loudness due to a grooved tongue channeling the airstream.
⭐ Key Takeaways
The most critical point is that speech sounds are best understood as articulatory gestures — dynamic movements with intrinsic timing — rather than static points or manners. The traditional terms for places of articulation should be thought of as articulatory targets, which may vary across languages. Students must recognize that languages differ greatly in their stop inventories (e.g., aspirated, implosive, ejective, murmured stops in Sindhi, Hausa, and Navajo) and that nasals can be voiceless in some languages despite being predominantly voiced in English. Finally, fricatives can be subdivided auditorily into sibilants ([s, z, ʃ, ʒ]) which have more acoustic energy and higher pitch, and non-sibilants ([f, θ, v, ð]).
🧠 Quick Revision Questions
- What does it mean to treat speech sounds as "gestures" rather than as exact articulatory locations?
- Give an example of a non-English bilabial fricative, and name a language that uses it.
- List three types of stop sounds found in Sindhi that are not found in English.
- How are voiceless nasals symbolized in phonetic transcription? Provide an example.
- What auditory feature distinguishes sibilant fricatives like [s, z, ʃ, ʒ] from non-sibilant fricatives like [f, θ]?
📘 Lecture 22 — Consonantal Gestures-II
📖 Overview: This lecture continues the exploration of articulatory gestures for English consonants, focusing on trills, taps, flaps, and laterals. It also provides a comprehensive summary of how different consonant sounds are classified based on shared articulatory features, emphasizing the gestural approach in phonology.
🗂️ Topics Covered
The lecture covers trills, taps, and flaps as approximant sounds found in world languages, including their varying articulatory gestures and lengths. It then discusses laterals as a distinct articulatory gesture, with special attention to the English lateral /l/. Finally, it summarizes the articulatory gestures for consonants, explaining how sounds are categorized into classes based on shared features and gestures, including primary and secondary articulation.
📝 Lecture Summary
Topic-110: Trills, Taps and Flaps
In approximants, trills, taps, and flaps are commonly found with different articulatory gestures in the world languages. These languages vary not only in the nature of the sounds (such as making it a usual /r/ approximant or unusual rhotic approximate [ɹ] found in American English) but also in the length of the sound (some making it a short trill, others a long one). The following table covers various types of these approximant sounds found in the world languages:
🔑 Definition — Trill: An articulatory gesture where one articulator (e.g., the tongue tip or uvula) is set into vibration by the airstream, producing a rapid series of taps.
🔑 Definition — Tap: A single, quick contact between an articulator (e.g., the tongue tip) and the roof of the mouth, with the movement being a single ballistic gesture.
🔑 Definition — Flap: A sound similar to a tap, but the articulator (e.g., the tongue tip) strikes the roof of the mouth in a downward (retroflex) motion.
| Type | IPA Symbol | Language Example |
|---|---|---|
| Voiced alveolar trill | r | Spanish |
| Voiced alveolar tap | ɾ | Spanish |
| Voiced retroflex flap | ɽ | Hausa |
| Voiced alveolar approximant | ɹ | English |
| Voiced alveolar fricative trill | ɻ | Czech |
| Voiced uvular trill | ʀ | French |
| Voiced uvular fricative or approximant | ʁ | Parisian French |
| Voiced bilabial trill | ʙ | Kele |
| Voiced labiodental flap | * | Margi |
💡 Why this matters: Understanding these different approximant gestures is crucial for accurately producing and transcribing sounds from languages other than English, particularly for field linguists and phoneticians.
Topic-111: Laterals
As an important articulatory gesture, the central–lateral opposition can be applied to all these manners of articulation, producing a lateral stop and a lateral fricative as well as a lateral approximant, which is by far the most common form of lateral sound. The only English lateral phoneme, at least in British English, is /l/ with allophones [l] as in led [lɛd] and [ɫ] as in bell [bɛɫ]. In most forms of American English, initial [l] has more velarization than is typically heard in British English initial [l]. In all forms of English, the air flows freely without audible friction, making this sound a voiced alveolar lateral approximant. It may be compared with the sound [ɹ] in red [ɹɛd], which is for many people a voiced alveolar central approximant. Laterals are usually presumed to be voiced approximants unless a specific statement to the contrary is made.
🔑 Definition — Central-lateral opposition: The distinction between sounds where air flows over the center of the tongue (central) versus sounds where air flows around the sides of the tongue (lateral).
🔑 Definition — Lateral approximant: A sound produced by placing the tongue tip on the alveolar ridge while allowing air to escape freely around one or both sides of the tongue.
📌 Example: The English /l/ has two allophones:
- Clear [l] — in initial position, e.g., led [lɛd]
- Dark [ɫ] — in final position, e.g., bell [bɛɫ]
Topic-112: Summary of the Articulatory Gestures
A number of sounds share various features of speech production in which the underlying units are represented by classes of functionally equivalent movement patterns (gestures). A particular gesture gradually increases its influence on the shape of the vocal tract, thus making many sounds similar in terms of articulatory features. Speech segments are modeled as sets of gestures that have their own intrinsic temporal structure, allowing them to overlap one way or the other. In the sections on various stop sounds in world languages, we studied 13 types of stop sounds (i.e., b, p, pʰ, bʱ, ɓ, b̰, kʼ, dn, nd, tɬ, tɬʼ, ts, tsʼ) and nine types of trill, tap and flap (together one category of approximants) sounds (i.e., r, ɾ, ɽ, ɹ, ɻ, ʀ, ʁ, ʙ, *) with a range of similarities in terms of their articulatory gestures. These examples show that the articulatory gestures for consonants have a wide range of possibilities of similar sounds in world languages.
🔑 Definition — Gesture: A functionally equivalent movement pattern used in speech production, with its own intrinsic temporal structure.
Topic-113: Explaining the Summary of the Articulatory Gestures
Based on a matrix of features specifying a particular characteristic of a segment and relating that particular characteristic to other similar sounds has made possible a range of similar sounds to be classified in the same class of segments. For example, an ‘oral gesture’ would specify all supraglottal characteristics (such as place and manner of articulation), and a ‘laryngeal gesture’ would specify characteristics of phonation. The notion is particularly used in dependency phonology, where ‘categorical’, ‘articulatory’ and ‘initiatory’ gestures are distinguished. Gestures, in turn, are analyzed into subgestures; for example, the initiatory gesture is analyzed into the subgestures of glottal stricture, airstream direction, and airstream source.
While discussing the articulatory gestures for consonants, it may be remembered that the case of consonants is even more complicated (than vowels). But still, a wide range of sounds are put together to classify them mainly on the basis of the characteristics of their primary articulation. However, sounds may also be taken together on the basis of their secondary articulation (e.g., added lip rounding).
🔑 Definition — Dependency phonology: A phonological theory that uses gestures as fundamental units, distinguishing categorical, articulatory, and initiatory gestures.
⭐ Key Takeaways
Trills, taps, and flaps are types of approximants with distinct articulatory gestures, and the IPA provides specific symbols for their transcription in world languages like Spanish, French, and Hausa. The English lateral approximant /l/ has two allophones ([l] and [ɫ]) depending on position, with American English showing more velarization in initial position. A gesture is a functionally equivalent movement pattern, and consonants can be classified by oral, laryngeal, and initiatory gestures. The 13 stop types and 9 trill/tap/flap types demonstrate the wide range of articulatory possibilities. Both primary articulation (e.g., place and manner) and secondary articulation (e.g., lip rounding) are used to group similar consonant sounds.
🧠 Quick Revision Questions
- What is the difference between a trill and a tap in terms of articulatory gesture?
- What are the two allophones of English /l/, and in which positions do they occur?
- How does American English initial [l] differ from British English initial [l]?
- What three types of gestures are distinguished in dependency phonology?
- Name three subgestures into which the initiatory gesture is analyzed.