eng507 — Final Term Summary (Lectures 23–42)
📘 Lecture 23 — Acoustic Phonetics-I
📖 Overview: This lecture introduces the field of acoustic phonetics, which is the scientific study of the physical properties of speech sounds as they travel from speaker to listener. It explains how speech can be analyzed objectively using computer software, and introduces the source-filter theory of speech production, which is essential for understanding how vowels are produced and distinguished.
🗂️ Topics Covered
The lecture covers the definition and explanation of acoustic phonetics as a scientific, objective method for analyzing speech sounds using physics and computer software. It details acoustic analysis methods for investigating durational and articulatory properties, with a specific focus on the acoustic analysis of vowels through formant frequencies. The source-filter theory of speech production is introduced and explained as a model where the larynx acts as a source of sound and the vocal tract acts as a filter, shaping the timbre of vowels and other speech sounds.
📝 Lecture Summary
Topic-114: Acoustic Phonetics
Acoustic phonetics is the study of the physics of the speech signal as a waveform traveling through air from the speaker’s mouth to the hearer’s ear via vibrations. These vibrations (physical properties) can be measured and analyzed using mathematical techniques and specially-developed computer software to produce spectrograms of speech. This field studies the relationship between vocal tract activity and resulting sounds, involving physics, computers, statistics, and lab experiments. Acoustic analysis is considered more objective and scientific than traditional auditory methods which rely on the trained human ear. It uses advanced speech software to analyze sound differences in pitch, loudness, and quality by providing detailed composition of energy, such as frequency on a spectrum.
🔑 Definition — Acoustic Phonetics: The branch of phonetics that studies the physical properties of speech sound as transmitted between the mouth of a speaker and the ear of a listener.
Topic-115: Explaining Acoustic Phonetics
Acoustic phonetics is a branch of phonetics devoted to the principles of physics and the latest technology, such as computer software like Praat. It is the scientific study of speech sounds wholly dependent on instrumental, lab-based techniques of investigation, involving electronics and physics (with mathematics as a prerequisite for advanced study). This course will cover basic features of speech sounds (consonants and vowels), fundamental experiments like recording and annotating speech, and distinguish among various forms and features of speech sounds. It also introduces the software Praat and its application in the phonetic analysis of human speech, with the goal of enabling learners to plan advanced-level applications and experiments.
💡 Why this matters: This explains that acoustic phonetics is not just a theoretical concept but a practical, lab-based science using specific tools like Praat to get objective data on speech.
Topic-116: Acoustic Analysis
According to phoneticians, acoustic analysis provides a clear, objective datum for investigating speech—the physical ‘facts’ of utterance. Acoustic evidence is often used to support analyses made in articulatory or auditory phonetics. However, one should not rely too heavily on acoustic analyses, as they can have mechanical limitations (e.g., the need to calibrate measuring devices accurately) and are often open to multiple interpretations. Acoustic analysis provides features of a sound and also tells us about the duration or length of a speech sound. Such analysis requires careful knowledge of the recording material and procedure. Thus, acoustic analysis describes durational characteristics, articulatory properties, and phonetic differences through physiological measurement.
🔑 Definition — Acoustic Analysis: The process of investigating speech to provide objective data on the physical properties of utterances, including durational characteristics and articulatory properties.
Topic-117: Acoustic Analysis of Vowels
Phonetic experts are particularly interested in analyzing vowels acoustically. They describe vowels in terms of numbers (how many vowels are possible in a language) and measure the actual frequencies of the formants (the formant structure of a language's vowels). Once these formants are taken, they can be represented graphically on a chart. The lecture references a figure from the textbook that gives the average values from several authorities for the frequencies of the first three formants in eight American English vowels. These vowels come from the words: heed, hid, head, had, hod, hawed, hood, who’d. At the end of the course, after Praat sessions, students will be able to record their own vowels and compare them to these values.
🔑 Definition — Formant: An overtone pitch that gives a vowel its distinctive quality. The lowest three formants distinguish vowels from each other.
Topic-118: Source Filter Theory of Speech Production
The source-filter theory is an important model of speech (e.g., vowel) production in acoustic phonetics. In this model, the source refers to the waveform of the vibrating larynx. Its spectrum is rich in harmonics, which gradually decrease in amplitude as their frequency increases. The filter is the various resonance chambers of the vocal tract (especially the movements of the tongue and lips), which act on the laryngeal source, reinforcing certain harmonics relative to others. The combination of these two elements (larynx as source, cavity as filter) is known as the source-filter model of speech production. Understanding this helps analyze changes in pitch, loudness, and quality. The quality of a vowel depends on its overtone structure (formants), meaning a sound contains a number of different pitches simultaneously: the pitch at which it is spoken and the various overtone pitches that give it its distinctive quality.
🔑 Definition — Source-Filter Theory: A model of speech production where the larynx serves as the source of sound energy (a waveform with harmonics), and the vocal tract acts as a filter that shapes that sound by reinforcing or damping specific frequencies to create different speech sounds.
📐 Formula/Model: Source (larynx vibration producing harmonics) + Filter (vocal tract resonating chambers) → Speech Sound (e.g., vowel with specific formants)
Topic-119: Explaining Source – Filter Mechanism
In this theory, the vocal tract is represented using a source-filter model, an idea used to synthesize speech. Sound travels from a noise-making source (vocal fold vibration) to the lips. At the lips, most sound energy radiates away for a listener to hear, while some sound energy reflects back into the vocal tract. The addition of reflected sound energy with the source energy tends to amplify energy at some frequencies and dampen it at others, depending on the length and shape of the vocal tract. The vocal folds (at the larynx) are the source of sound energy, and the cavity (vocal tract, due to the interaction of reflected sound waves) is a frequency filter altering the timbre of the vocal fold sound. This mechanism is also at work in musical instruments; for example, in brass instruments, the vibrating lips in the mouthpiece are the noise source, and the long brass tube provides the filter.
💡 Why this matters: This explains the physical mechanism of how the source (vocal folds) and filter (vocal tract) interact to produce the distinctive formants of vowels, making the abstract theory concrete by comparing it to musical instruments.
⭐ Key Takeaways
For the exam, you must remember that acoustic phonetics is the scientific, objective study of the physical properties of speech sounds using instrumental techniques and software like Praat. You need to understand how acoustic analysis provides objective data on durational and articulatory properties of sounds, with a particular focus on vowel analysis using formant frequencies (F1, F2, F3) which distinguish vowels from each other. The most critical concept to master is the source-filter theory, which models speech production as a source (the vibrating larynx producing harmonics) and a filter (the vocal tract that reinforces or dampens these harmonics). Finally, remember that the interaction of reflected sound waves in the vocal tract amplifies some frequencies and dampens others based on the tract's length and shape, and this same mechanism works in musical instruments like brass instruments.
🧠 Quick Revision Questions
- What makes acoustic phonetics more "objective and scientific" than traditional auditory methods of speech analysis?
- What are formants, and how many of them are typically used to distinguish one vowel from another?
- According to the source-filter theory, what is the "source" and what is the "filter" in speech production?
- What happens to sound energy when it reaches the lips in the source-filter mechanism, and how does this affect the sound wave?
- What is the name of the computer software mentioned in this lecture that is used for phonetic analysis?
📘 Lecture 24 — Acoustic Phonetics-II
📖 Overview: This lecture explores the acoustic properties of speech sounds, focusing on how the vocal tract shapes sound through tube models and perturbation theory. It explains the relationship between vocal tract configurations and formant frequencies for vowels, then introduces the more complex acoustic features of consonants.
🗂️ Topics Covered
This lecture covers tube models of the vocal tract and their role in understanding vowel resonance, perturbation theory for predicting formant frequency changes through constrictions, acoustic analysis of vowels using spectrograms with F1 and F2 formant relationships, and the acoustic properties of consonants including stops, nasals, approximants, and glides.
📝 Lecture Summary
Topic-120: Tube Models
To understand vowels and their formants, we use the concept of the vocal tract as a tube. The formants that characterize different vowels result from different shapes of the vocal tract. Air in the vocal tract vibrates depending on its size and shape, set in motion by the vocal folds. Every time the vocal folds open and close, there is a pulse of acoustic energy. Regardless of the rate of vibration at the source (vocal folds), the air in the vocal tract (the filter) will resonate at specific frequencies as long as the position of the vocal organs remains the same. Due to the complex shape of the filter, the air vibrates in more than one way at once. In most voiced sounds, three formants are produced each time the vocal folds vibrate. Crucially, resonance in the filter is independent of the rate of vibration of the source.
💡 Why this matters: This demonstrates the source-filter theory — the source (vocal folds) determines pitch, while the filter (vocal tract) determines the vowel quality through formant frequencies.
Topic-121: Explaining the Tube Models
This is an old idea from Hermann Helmholtz that a vowel is the rapid repetition of its peculiar two or three notes (the first two or three formants). All voiced sounds are distinguishable by their formant structure (frequencies). When vocal fold pulses are produced at a steady rate, the utterance is on a monotone. What we hear as changes in pitch are actually changes in the overtones of this monotone voice. These overtone pitch variations convey the quality of voiced sounds. The rhythm of a sentence is apparent because overtone pitches occur only when the vocal folds vibrate. This tube model provides understanding of resonance, resonant frequencies, the source-filter model, and the relationship between articulatory configurations and acoustic consequences.
Topic-122: Perturbation Theory
A vocal tract modeled as a tube with uniform diameter has simultaneous resonance frequencies. These resonance frequencies change predictably when the tube is squeezed at various locations. We can model vowel acoustics in terms of perturbations of the uniform tube. For example, when lips are rounded, the diameter at the lips is smaller than elsewhere. Perturbation theory says that with the acoustic effect of constriction at the lips, we can predict formant frequency differences between rounded and unrounded vowels. For each formant, there are locations where constriction will cause the formant frequency to rise, and locations where constriction will cause it to fall.
🔑 Definition — Perturbation Theory: A theory describing how constrictions at specific locations in the vocal tract cause predictable changes (increase or decrease) in resonance frequencies (formants), depending on whether the constriction is at a node or anti-node of that resonance mode.
Topic-123: Explaining Perturbation Theory
According to perturbation theory, resonance occurs in a uniform tube where one end is closed and the other end is open (the source-filter idea). The theory tells whether each resonance frequency increases or decreases when a small modification occurs in the diameter at a local region of the tube.
The rule:
- The resonance frequency decreases when a constriction is located at an anti-node of that resonance mode.
- The resonance frequency increases when a constriction is located at a node of that resonance mode.
Thus, resonance frequencies change according to the position and nature of modifications in the vocal tract.
📌 Example: If a constriction (like lip rounding) occurs at an anti-node of the first formant, F1 will decrease. If the same constriction occurs at a node of the second formant, F2 will increase.
Topic-124: Explaining Acoustic Analysis (Vowels)
Using computer programs, vowel sounds can be analyzed through spectrograms. In spectrograms, time runs left to right, frequency is shown on the vertical scale, and intensity of each component is shown by degree of darkness. Dark bands show concentrations of energy at particular frequencies, revealing source and filter characteristics. Traditional articulatory descriptions of vowels are related to formant frequencies. The first two formants are most important:
- First formant (F1) is inversely related to vowel height: higher vowels have lower F1, lower vowels have higher F1.
- Second formant (F2) is related to vowel frontness: front vowels have higher F2, back vowels have lower F2.
When F1 and F2 are plotted, the vowel chart structure closely matches traditional articulatory descriptions.
📐 Formula: F1 ∝ 1/vowel height (higher vowel → lower F1) 📐 Formula: F2 ∝ vowel frontness (fronter vowel → higher F2)
Topic-125: Acoustics of Consonants
The acoustic properties of consonants are usually more complicated than those of vowels. A consonant can be described as a particular way of beginning or ending a vowel sound because during consonant production there is no distinguishing feature prominently visible. During closures of voiced stops [b, d, g] there is virtually no difference, and absolutely none during voiceless stops [p, t, k] because there is only silence. Each stop sound conveys its quality by its effect on the adjacent vowel. While some consonants have vowel-like structures (nasals, approximants, glides), most consonants have totally different acoustic features. For stops, the complete closure is easily visible on a spectrogram as a simple thin line.
To analyze consonant acoustic features, use instructions from Chapter 8 of the course book. Particular sound combinations are important as neighboring vowels help in guessing exact consonant sounds.
⭐ Key Takeaways
The most critical concepts from this lecture are: (1) The vocal tract acts as a tube filter whose shape determines formant frequencies independently of vocal fold vibration rate, (2) Perturbation theory predicts that constrictions at nodes increase formant frequencies while constrictions at anti-nodes decrease them, (3) For vowels, F1 inversely relates to vowel height and F2 relates to vowel frontness, as visible on spectrograms, (4) Consonants are acoustically more complex than vowels, with stops showing silence during closure and their identity revealed through transitions on adjacent vowels, and (5) The source-filter model explains how the glottal source provides the fundamental frequency while the vocal tract filter shapes the acoustic output into distinct speech sounds.
🧠 Quick Revision Questions
- What is the relationship between vocal fold vibration rate and formant frequencies?
- According to perturbation theory, what happens to a formant frequency when a constriction occurs at a node of that resonance mode?
- How are F1 and F2 related to vowel height and frontness respectively?
- Why are the acoustic properties of consonants generally more complicated than those of vowels?
- How do stops like [p, t, k] convey their identity on a spectrogram if the closure period shows only silence?
📘 Lecture 25 — Acoustic Phonetics-III
📖 Overview: This lecture focuses on interpreting spectrograms of speech sounds, providing specific acoustic cues for consonants and techniques for analyzing connected speech. It also addresses individual differences in speech production and how to analyze these variations for applications in forensic linguistics and speech recognition technology.
🗂️ Topics Covered
This lecture covers five main topics: explaining the acoustics of consonants with specific spectrographic cues for place and manner of articulation, the importance and practical reasons for learning spectrogram interpretation, useful techniques for interpreting spectrograms in connected speech, an introduction to individual differences in speech acoustics, and methods for analyzing and comparing individual speakers' acoustic data.
📝 Lecture Summary
Topic-126: Explaining the Acoustics of Consonants
This section provides a reference list of spectrographic features for consonant sounds. Voiced consonants show vertical striations from vocal fold vibration. Bilabial sounds have a low locus for both F2 and F3. Alveolar sounds have an F2 locus around 1700–1800 Hz. Velar sounds typically have a high F2 locus. Retroflex sounds show a general lowering of F3 and F4. Stops show a gap in the pattern, with a burst for voiceless stops and a sharp formant beginning for voiced stops. Fricatives show random noise patterns in higher frequency regions. Nasals have a formant structure similar to vowels, with formants at 250, 2500, and 3250 Hz. Laterals also have a vowel-like formant structure, with formants at 250, 1200, and 2400 Hz. Approximants have a formant structure similar to vowels, usually changing.
🔑 Definition — Locus: The typical frequency region where formants begin for a consonant, especially relevant for stop consonants.
Topic-127: Interpreting Spectrograms
This section explains why learning spectrogram interpretation is valuable. Words recorded in isolation are easier to interpret than connected speech. Spectrograms are useful for researchers working on the nature of sounds (e.g., phoneme vs. allophone). They increase our understanding of speech sounds and their behavior in different forms. Practice on spectrograms teaches the characteristics of speech sounds. Spectrograms are important for experts in speech signal processing and are part of techniques used in speech recognition. They enable exploration of the complex nature of speech structure and are used in machine translation.
💡 Why this matters: Learning spectrogram interpretation is not just an academic exercise—it has direct applications in speech technology, linguistics research, and automated systems.
Topic-128: Useful Techniques for Interpreting Spectrograms
This section provides practical strategies for interpreting spectrograms, especially for connected speech. Start by analyzing sounds one by one, keeping their individual class characteristics in mind. Carefully observe the overall structure, especially the frequency scale. When interpreting consonants, also analyze the behavior of adjacent vowels. Pay more attention to the first two formants (especially for vowels). Watch for a burst and aspiration in stop sounds. Remember that in vowels, F1 is inversely related to height (lower F1 = higher vowel), and F2 is related to backness. It is also possible to observe manner changes, such as a stop weakening to a fricative or approximant, affrication of a stop, and separation of trills from taps and flaps.
📐 Formula: F1 ∝ 1/vowel height → As vowel height increases, F1 decreases. 📐 Formula: F2 ∝ backness → F2 varies with the front-back position of the tongue.
Topic-129: Individual Differences
This section introduces the concept of individual variation in speech acoustics. At times, we need to know about idiosyncratic features of speech to distinguish them from a typical speech community or for forensic investigations. Spectrograms show relative quality (e.g., one speaker may have a higher level of vowels than another). Formant plots of an average speaker can be compared with those of a particular speaker. When two speakers record the same vowel qualities, their relative positions on a formant chart will be similar, but the absolute values of formant frequencies will differ. A key strategy is to use the values of F4, which may be an indicator of individual head size.
🔑 Definition — Idiosyncratic features: Individual-specific speech characteristics that help distinguish one speaker from another, relevant in forensic phonetics.
Topic-130: Analyzing Individual Differences
This section explores methods for comparing speakers. As an alternative method, a speaker's complete vowel range may be considered representative of their personal features, which can then be compared with the formant frequency of each vowel relative to that speaker's total formant range. Phoneticians are still working on comparing acoustic data between individuals and improving speech recognition systems. Experts in applied phonetics and computer speech technology are exploring areas like intonation and rhythm. For correct estimates, researchers consider working with syntax, pragmatics, rhythmic structure, and individual segmental influences together to develop programs for speech recognition. Speech identification is also directly needed by legal experts, and voice-prints are now largely used by lawyers. Spectrograms are a potential area for future research studies.
💡 Why this matters: Understanding individual differences is crucial for forensic linguistics, speaker identification, and developing more robust speech recognition systems that can handle different voices.
⭐ Key Takeaways
The critical points to remember from this lecture are: First, each consonant class has specific spectrographic cues (e.g., fricatives show random noise, stops show gaps with bursts, nasals have formants at 250, 2500, 3250 Hz). Second, spectrogram interpretation is valuable for linguistics research, speech recognition, and machine translation. Third, when analyzing spectrograms, start with individual sounds, watch for the frequency scale, and pay attention to the first two formants (F1 for vowel height, F2 for backness). Fourth, individual differences exist in absolute formant values even when relative vowel positions are similar, and F4 can be a useful indicator of head size. Finally, spectrogram analysis is directly relevant in forensic linguistics for speaker identification and in developing advanced speech technology systems.
🧠 Quick Revision Questions
- What are the spectrographic cues for identifying a velar stop consonant, especially regarding the second formant?
- Why is F1 inversely related to vowel height, and what does this mean for interpreting a high vowel like [i]?
- What are three practical reasons for learning to interpret spectrograms in phonetics and phonology?
- How can F4 be used to analyze individual differences between speakers?
- What is the difference between analyzing a word recorded in isolation versus analyzing connected speech on a spectrogram?
📘 Lecture 26 — Vowels and Vowel-Like Articulations-I
📖 Overview: This lecture introduces the fundamental phonetic features that distinguish vowels from consonants and describes how vowels are classified. It explains the cardinal vowel system developed by Daniel Jones, including both primary and secondary cardinal vowels, providing a universal reference framework for describing vowel sounds across all languages.
🗂️ Topics Covered
This lecture covers the definition and phonetic features of vowels (lip-rounding, tongue part, and tongue height), the concept of cardinal vowels as a reference system, the secondary cardinal vowels used for non-European languages, and a direct comparison between primary and secondary cardinal vowel systems using numerical codes and articulatory features.
📝 Lecture Summary
Vowels and Vowel-like Articulation
The fundamental distinction between consonants and vowels is that vowels make the least obstruction to the flow of air. Vowels are almost always found at the center of a syllable, and it is very rare to find any sound other than a vowel that can stand alone as a whole syllable. Phonetically, each vowel has several features that distinguish it from other vowels. These include: lip-rounding, which can be rounded (for sounds like /u:/), neutral (as for schwa /ə/), or spread (as in /i:/ in "cheese" /tʃi:z/); the part of the tongue (front, middle, or back) that may be raised—compare /æ/ as a front vowel with /ɑ:/ as a back vowel; and tongue height (and jaw position), where the tongue may be raised ‘close’ to the roof of the mouth (for close vowels like /i:/ or /u:/) or left ‘low’ in the mouth with the jaw comparatively ‘open’ (for open vowels like /a:/ and /æ/). In British phonetics, terms like ‘close’ and ‘open’ are used, whereas in American phonetics ‘high’ and ‘low’ are used. Generally, these three aspects are described for vowels: lip-rounding, part of the tongue, and tongue height. Additional characteristics like nasality (whether a vowel is nasal or not) are also used in various languages of the world.
Cardinal Vowels
To classify vowels independently of a particular language, British phonetician Daniel Jones introduced a system in the early 20th century based on a set of “cardinal vowels” comprising eight vowels used as reference points (like the corners and sides of a map). Jones was strongly influenced by French phonetician Paul Passy, and the set is claimed to be similar to the vowels of educated Parisian French of his time. The cardinal vowel system is a chart or four-sided figure (the exact shape has changed over time) with eight corners, as seen on the IPA chart. It is a diagram used for both rounded and unrounded vowels. Jones proposed a primary and a secondary set of cardinal vowels. The primary set includes eight vowels (1 to 8): front unrounded vowels [i, e, ε, a], the back unrounded vowel [ɑ], and the rounded back vowels [ɔ, o, u].
Secondary Cardinal Vowels
To explain vowels of non-European languages, a set of secondary cardinal vowels was introduced by Daniel Jones as a precise set of references. The main difference between primary and secondary cardinal vowels is related to lip-rounding, as in some languages, lip-rounding is possible for front vowels. By reversing the lip position compared to primary cardinal vowels, the secondary series is produced (e.g., rounding the lips for front vowels). The secondary cardinal vowels (with numerical codes and features) are:
- 9. Close (high) front rounded vowel [y]
-
- Close-mid front rounded vowel [ø]
-
- Open-mid front rounded vowel [œ]
-
- Open (low) front rounded vowel [ɶ]
-
- Open (low) back rounded vowel [ɒ]
-
- Open-mid back unrounded vowel [ʌ]
-
- Close-mid back unrounded vowel [ɤ]
-
- Close (high) back unrounded vowel [ɯ]
Two more vowels, 17 and 18, represented by symbols [ɨ] and [ʉ] respectively, represent the highest possible position at the center of the tongue.
💡 Why this matters: The secondary cardinal vowel system extends the primary system to cover sounds found in languages like French, German, and Turkish, where front rounded vowels (like [y] and [ø]) are common.
Comparing Primary VS Secondary Cardinal Vowels
Secondary cardinal vowels are easy to understand in connection with the primary cardinal vowel system. Secondary cardinal vowels from 9 to 13 [y, ø, œ, ɶ, ɒ] are the rounded counterparts of primary cardinal vowels from 1 to 5 [i, e, ε, a, ɑ]. Similarly, secondary cardinal vowels from 14 to 16 [ʌ, ɤ, ɯ] are the unrounded equivalents of primary cardinal vowels from 6 to 8 [ɔ, o, u] respectively. Two further cardinal vowels (17 and 18, symbolized by [ɨ] and [ʉ]) represent the highest point at the center the tongue can reach. The entire vowel system of human languages is usually shown in the form of the cardinal vowel diagram (or cardinal vowel quadrilateral), which resembles the human tongue and is divided into eight corners. The aim is to give an approximate configuration of the degree and direction of tongue movement involved in vowel production. These diagrams are successfully used for describing vowel systems in dialects and languages of the world.
🔑 Definition — [Cardinal Vowels]: A system of eight reference vowels (later extended to 18) introduced by Daniel Jones, independent of any particular language, used as fixed points to describe and classify all vowel sounds.
📐 Formula:
- Primary → Secondary relationship:
- Front vowels (1-5) + Rounding → Secondary front vowels (9-13)
- Back vowels (6-8) + Unrounding → Secondary back vowels (14-16)
- Numerical mapping:
- 1[i] ↔ 9[y] (rounding added)
- 5[ɑ] ↔ 13[ɒ] (rounding added)
- 6[ɔ] ↔ 14[ʌ] (unrounding added)
- 8[u] ↔ 16[ɯ] (unrounding added)
📌 Example: Compare primary vowel 1 [i] (close front unrounded, as in English “see”) with secondary vowel 9 [y] (close front rounded, as in French “tu” meaning “you”). The difference is solely lip-rounding: both have the tongue in the same high front position, but [i] has spread lips while [y] has rounded lips.
⭐ Key Takeaways
The fundamental features distinguishing vowels are lip-rounding (rounded, neutral, or spread), part of the tongue (front, central, or back), and tongue height (close/open or high/low). Daniel Jones’ cardinal vowel system provides a universal reference framework with 8 primary and 10 secondary vowels, where the secondary set is derived by reversing the lip position of the primary set. Secondary cardinal vowels 9-13 are rounded counterparts of primary vowels 1-5, while vowels 14-16 are unrounded counterparts of primary vowels 6-8, with vowels 17 and 18 representing the highest central tongue position. The cardinal vowel diagram (quadrilateral) maps tongue movement and allows systematic description of all vowel sounds across languages.
🧠 Quick Revision Questions
- What are the three main phonetic features used to describe vowels?
- How many primary cardinal vowels are there, and what are their basic articulatory characteristics?
- What is the main difference between primary and secondary cardinal vowels?
- Which secondary cardinal vowel corresponds to primary cardinal vowel 3 [ε], and what feature changes?
- What do cardinal vowels 17 [ɨ] and 18 [ʉ] represent in terms of tongue position?
📘 Lecture 27 — Vowels and Vowel-like Articulations-II
📖 Overview: This lecture explores vowel variation across different accents of English and other world languages. It examines the vowel systems of BBC English, American English varieties, and languages such as Spanish and Japanese. The lecture also introduces the concepts of Advanced Tongue Root (ATR), a key feature in African language phonology, and rhotacization, which distinguishes rhotic from non-rhotic English accents.
🗂️ Topics Covered
The lecture begins by examining vowel differences across various accents of English, focusing on American varieties and BBC English. It then surveys vowel systems in languages like Spanish, Japanese, Danish, Urdu, and Punjabi. Special attention is given to Advanced Tongue Root (ATR) as a phonological feature in Akan and other African languages, and finally, rhotacized vowels are explained as a distinguishing feature between rhotic and non-rhotic English accents.
📝 Lecture Summary
Topic-135: Vowels in Other Accents of English
Vowel sounds in various accents of English provide a solid basis for comparisons and contrasts. Researchers have explored this aspect in various Englishes and published descriptions of the auditory quality of vowels. Spectral structure, by referring to the average formant frequencies of vowel systems, provides interesting comparisons. For example, the accent of American (newscasters) English has a fairly conservative difference (in the first two formants) compared with Californian English. The Californian English does not maintain a contrast between the vowels in cot and caught (both are spoken as the same [ɑ]). Moreover, Californians have a higher vowel (lower first formant) in [eɪ] than in [ɪ]. Their high back vowels seem more front because they have a higher second formant. Additionally, vowel /ʊ/ is often pronounced with spread lips in this variety. In a number of northern cities in the United States (e.g., Pittsburgh and Detroit), [æ] is spoken very close to [ɛ] (as raised with decreased F1).
Topic-136: Vowels in BBC English
The vowels of BBC accent differ from American English in both number (20 in total – both pure vowels and diphthongs) and quality. British English speakers distinguish the vowel [ʌ] in cut from the vowel [ɜ] in curt. BBC English does not have any r-coloring (rhotacization). The main feature of BBC English is the distinction between three back vowels: [a] as in father and cart, [ɑ] as in bother and cot, and [ɔ] as in author and caught. BBC vowels are interesting in making it a standard variety of English.
Topic-137: Vowels in Other Languages
Vowels in languages other than English vary not only in number but also in quality and features. Spanish has a simple system contrasting only five vowel sounds [i, e, a, o, u], though these symbols do not have the same values as English cardinal vowels. Japanese also has a set of five vowels [i, e, a, o, u] but very different in a narrower transcription than Spanish. Danish contrasts three front rounded vowels which may occur in long or short forms. Focusing on Pakistani regional languages, Urdu and Punjabi have very different nasal vowels (nasality is phonemic). Urdu has seven nasal vowels such as /he/ vs. /hẽ/ or /hi:/ vs. /hĩ:/ as in word /nahĩ:/ meaning ‘no’. Nasalization in vowels is a common feature of Indo Aryan languages.
Topic-138: Advanced Tongue Root (ATR)
Normally, vowel description uses three features: height of tongue, backness of tongue, and lip rounding. However, this does not cover all variation. One such variation is Advanced Tongue Root (ATR), found in Akan language spoken in Ghana. Vowels produced with ATR involve the furthest-back part of the tongue, opposite to the pharyngeal wall, which is not normally involved in speech sound production — also called the radix (articulations of this type may be described as radical). ATR is a kind of articulation where movement of the root of the tongue expands the front–back diameter of the pharynx. It is used phonologically in Akan (and some other African languages) as a factor in vowel harmony. The opposite direction of movement is called retracted tongue root (RTR). ATR is related to the size of the pharynx, creating comparatively large (+ATR: root forward and larynx lowered) and small pharyngeal cavity (-ATR: no advanced tongue root). Akan contrasts between two sets of vowels: +ATR and -ATR.
🔑 Definition — Advanced Tongue Root (ATR): An articulation in which the root of the tongue moves forward, expanding the front–back diameter of the pharynx, used phonologically in Akan and other African languages for vowel harmony.
💡 Why this matters: ATR expands the traditional three-feature description of vowels (height, backness, rounding) and shows that cross-linguistic vowel systems can use pharyngeal cavity size as a contrastive feature.
Topic-139: Rhotacized Vowels
Rhotacization (or rhotacized vowel) is a term used in English phonology referring to dialects or accents where /r/ is pronounced following a vowel, as in words ‘car’ and ‘cart’. Varieties of English are divided on this basis: rhotic varieties have /r/ in all phonological contexts, while non-rhotic varieties (such as Received Pronunciation) have /r/ only before vowels as in ‘red’ and ‘around’. Vowels which occur after retroflex consonants are sometimes called rhotacized vowels. While BBC pronunciation is non-rhotic, many accents of the British Isles are rhotic, including most of the south and west of England, much of Wales, and all of Scotland and Ireland. Most American English speakers speak with a rhotic accent, but there are non-rhotic areas (e.g., the Boston area, lower-class New York, and the Deep South).
🔑 Definition — Rhotacization: The pronunciation of /r/ following a vowel; a feature that distinguishes rhotic dialects (where /r/ appears in all contexts) from non-rhotic dialects (where /r/ appears only before vowels).
⭐ Key Takeaways
This lecture demonstrates that vowel systems vary significantly across accents and languages in number, quality, and articulation features. BBC English maintains a three-way back vowel distinction ([a], [ɑ], [ɔ]) and is non-rhotic, while many American varieties merge certain vowels (e.g., cot/caught) and show regional differences in vowel raising. Languages like Spanish and Japanese have five-vowel systems, but their phonetic values differ from English cardinal vowels. Advanced Tongue Root (ATR) is a crucial feature in African languages like Akan, contrasting vowel sets through pharyngeal cavity size, and is related to vowel harmony. Rhotacization divides English accents into rhotic and non-rhotic types, with significant regional variation in both Britain and America.
🧠 Quick Revision Questions
- How does Californian English differ from conservative American English in terms of the cot/caught contrast and the vowel /ʊ/?
- What are the three back vowels distinguished in BBC English, and what example words illustrate each?
- How do the vowel systems of Spanish and Japanese differ in terms of phonetic transcription, even though both have five vowels?
- What is Advanced Tongue Root (ATR), and how does it function in the Akan language's vowel harmony system?
- What is the difference between rhotic and non-rhotic English accents, and give one example of a rhotic and one non-rhotic variety.
📘 Lecture 28 — VOWELS AND VOWEL-LIKE ARTICULATIONS-I II
📖 Overview: This lecture explores features of vowels beyond basic tongue height and backness, including nasalization, semivowels, and secondary articulatory gestures. Understanding these features is crucial for accurately describing and transcribing the vowel systems of languages like Urdu, Punjabi, and French, which differ significantly from English.
🗂️ Topics Covered
This lecture begins by defining vowel nasalization and its diacritic representation, then provides a summary of all six vowel quality features. It introduces semivowels as sounds that function like consonants but are phonetically similar to vowels. The lecture then explains four types of secondary articulatory gestures—palatalization, velarization, pharyngealization, and labialization—and concludes with a summary of these gestures for vowel quality.
📝 Lecture Summary
Topic-140: Nasalization in Vowels
Nasalization is a vowel feature not covered by the traditional method of description. In languages like Urdu, Punjabi, and French, nasal vowels contrast with oral vowels. For speakers of languages like English that lack nasal vowels, learning this feature requires pronouncing a vowel while keeping the soft palate lowered. Urdu has seven nasal vowels, such as /he/ (‘is’) vs. /hẽ/ (‘are’). Nasalization is a common feature of Indo Aryan languages.
🔑 Definition — Nasalization: The production of a vowel sound with air escaping through the nose, achieved by lowering the soft palate. 📐 Diacritic: [ ̃] (tilde) — placed above the phonetic symbol to indicate nasality. 📌 Example: In Urdu, /he/ (meaning ‘is’) is an oral vowel, while /hẽ/ (meaning ‘are’) is its nasal counterpart. The difference is solely the lowered soft palate.
Topic-141: Summary of Vowel Quality
This topic summarizes all six features of vowel quality. Two features—height and backness of the tongue—are used to contrast vowels in nearly all languages. Four other features are used less frequently:
- Lip-rounding
- Rhotacization
- Nasalization
- Advanced Tongue Root (ATR)
The lecture provides the acoustic correlates for each of these six features.
🔑 Acoustic Correlates of Vowel Features: 📐 Height: Frequency of Formant 1 (F1) — inversely related (higher vowel = lower F1) 📐 Backness: Difference between frequencies of F2 and F1 📐 Rhotacization: Frequency of Formant 3 (F3) 📐 Rounding: Lip position (rounded, half-rounded, or neutral) 📐 ATR: Width of the pharynx (ATR or Retracted Tongue Root - RTR) 📐 Nasalization: Position of the soft palate (lowered for nasal, raised for oral)
💡 Why this matters: These acoustic correlates are the physical measurements that allow phoneticians to objectively differentiate vowel qualities using spectrograms.
Topic-142: Semivowels
Semivowels (also called approximants) are a class of sounds that function like consonants but are phonetically similar to vowels. For example, in English, /w/ in ‘wet’ and /j/ in ‘yet’ are semivowels. When they occur at the onset of a syllable, they function as consonants. When pronounced slowly, /w/ resembles the vowel [u] and /j/ resembles the vowel [i]. French has three semivowels: /j/, /w/, and /ɥ/ (as in ‘huit’ /ɥit/, meaning ‘eight’). The IPA chart also lists a semivowel corresponding to the back close unrounded vowel /ɯ/.
🔑 Definition — Semivowel/Approximant: A sound that is phonetically like a vowel but functions phonologically as a consonant, typically occurring at the beginning of a syllable. 📌 Example: In English ‘wet’ /wet/, /w/ is a semivowel. In French ‘fruit’ /frɥi/, /ɥ/ is a semivowel.
Topic-143: Secondary Articulatory Gestures (SAG)
A secondary articulation is an articulatory gesture with a lesser degree of closure that occurs at approximately the same time as the primary gesture. This differs from co-articulation, where both gestures are of equal value. Four types of secondary articulations are associated with vowels:
- Palatalization: Adding a high front tongue gesture (like /i/)
- Velarization: Raising the back of the tongue
- Pharyngealization: Superimposing a narrowing of the pharynx/larynx
- Labialization: Adding lip-rounding
🔑 Definition — Secondary Articulation: A less constricted articulatory gesture that occurs simultaneously with a primary (more constricted) gesture in the vocal tract.
Topic-144: SAGs Discussion
This topic provides detailed descriptions of each secondary articulatory gesture.
Palatalization is the addition of a high front tongue gesture (like [i]) to the primary gesture. The diacritic is a small [ʲ] placed above the primary symbol. A palatalized consonant has a /j/-like quality. The term palatalization can also refer to a process where the place of articulation shifts towards the center of the hard palate.
🔑 Definition — Palatalization: Raising the front of the tongue close to the palate while making an articulatory closure at another point. 📐 Diacritic: [ʲ] 📌 Example: In many languages, /t/ can be palatalized to [tʲ], sounding like /tj/.
Velarization involves raising the back of the tongue, adding a [u]-like quality (without lip rounding). A classic English example is the velarized or dark /l/ ([l̴]) found at the end of syllables in words like kill, pill, sell, and will.
🔑 Definition — Velarization: Raising the back of the tongue towards the velum while making a primary articulation elsewhere. 📐 Diacritics: [ˠ] and [ ̴] 📌 Example: The /l/ in ‘pill’ is a velarized /l/ [l̴], whereas the /l/ in ‘lip’ is a non-velarized (clear) /l/ [l].
Pharyngealization is the superimposition of a narrowing of the pharynx.
🔑 Definition — Pharyngealization: Constructing the pharynx (the region above the larynx) while making the primary articulation. 📐 Diacritics: [ ̴] and [ˤ] 📌 Example: Pharyngealized consonants are common in Arabic and some Semitic languages.
Labialization is the addition of lip-rounding to another primary articulation.
🔑 Definition — Labialization: Adding lip rounding to a consonant or vowel. 📐 Diacritic: [ʷ] 📌 Example: Arabic /tʷ/ and /sʷ/ are labialized consonants. Nearly all kinds of consonants can be labialized.
Topic-145: Summary of the SAGs
This topic summarizes the four secondary articulatory gestures for vowel quality. Any speech sound may or may not have any of these four gestures. It is possible for a sound to be labialized and also have one of the other three secondary articulations simultaneously (e.g., [mʷ]).
The main features are:
- Palatalization [ʲ]: Raising of the front of the tongue (like /i/ vowel).
- Velarization [ˠ] and [ ̴]: Raising of the back of the tongue (like [u]-like sound).
- Pharyngealization [ˤ]: Retracting of the root of the tongue.
- Labialization [ʷ]: Rounding of the lips (e.g., Arabic [sʷ] and [tʷ]).
⭐ Key Takeaways
The most critical concepts from this lecture are the definitions and distinctions between nasalization, semivowels, and secondary articulations. Nasalization uses the soft palate to produce nasal or oral vowels. Semivowels (approximants) like /w/ and /j/ function as consonants but are phonetically vowel-like. The four secondary articulatory gestures—palatalization, velarization, pharyngealization, and labialization—each add a specific vowel-like quality to a primary consonant articulation, and are represented by distinct IPA diacritics. These features are essential for accurately characterizing the diverse vowel and consonant systems found across world languages.
🧠 Quick Revision Questions
- What is the IPA diacritic for nasalization, and what articulator is manipulated to produce it?
- List three languages that have contrastive nasal vowels.
- What is the difference between a semivowel and a vowel in terms of their function in a syllable?
- Name the four types of secondary articulatory gestures, and state which vowel sound each one resembles (e.g., /i/, /u/, etc.).
- What is the difference between a primary articulation and a secondary articulation? Give an example.
📘 Lecture 29 — Suprasegmental Features-I
📖 Overview: This lecture introduces suprasegmental features—vocal effects that extend beyond individual speech sounds. It focuses on defining and examining the syllable and stress as fundamental suprasegmental units, explaining their structures, types, and linguistic functions, including how stress can change meaning at word and sentence levels.
🗂️ Topics Covered
The lecture begins with an introduction to suprasegmental features, then defines and explains the syllable in detail, including its structure (onset, nucleus, coda, rhyme) and types. It then covers stress as a suprasegmental feature, exploring its types and categories (primary, secondary, tertiary), and concludes with the distinction between lexical stress and emphatic (sentence) stress.
📝 Lecture Summary
Topic-146: Introduction to Suprasegmental (SS) Features
Suprasegmental features are vocal effects (like tone, intonation, stress, etc.) that extend over more than one sound (segment) in an utterance. The term combines supra (meaning ‘above’ or ‘beyond’) and segments (which are sounds or phonemes). Major suprasegmental features include pitch, stress, tone, intonation, and juncture. These features are only meaningful when applied above the segmental level—i.e., on more than one segment. Phonological studies are divided into two fields: segmental phonology and suprasegmental phonology. Suprasegmental features have been extensively explored in recent decades, leading to many theories about their application and description.
Topic-147: Syllable
The syllable is a fundamentally important unit in both phonetics and phonology. Experts keep the phonetic notions of the syllable separate from the phonological ones. It is easy to understand but very difficult to define. Simply put, syllables are the parts of a word into which it can be divided, e.g., mi-ni-mi-za-tion or sup-re-seg-men-tal.
🔑 Definition — Syllable (phonetic) : From a speech production point of view, a syllable consists of a movement from a constricted or silent state to a vowel-like state and then back to a constricted or silent state.
Phonetically, the flow of speech typically alternates between vowel-like states (where the vocal tract is open and unobstructed) and consonant-like states (where there is some obstruction to airflow). Acoustically, this means the speech signal shows a series of peaks of energy (corresponding to vowel-like states) separated by troughs of lower energy (or sonority).
Topic-148: Explaining Syllable
Phonologists are interested in the structure of a syllable. It can be divided into three possible parts: the beginning (onset), the middle (nucleus or peak), and the end (coda). The combination of the nucleus (peak) and the coda is called the rhyme.
🔑 Definition — Nucleus: The central, mandatory part of a syllable, usually a vowel sound. A syllable must have a nucleus (at least one phoneme), while the onset and coda are optional.
🔑 Definition — Onset: The beginning sound(s) of a syllable, which are optional consonants.
🔑 Definition — Coda: The ending sound(s) of a syllable, which are optional consonants.
🔑 Definition — Rhyme: The combination of the nucleus and the coda within a syllable.
The study of the sequences of phonemes is called phonotactics, and it seems that the phonotactic possibilities of a language are determined by its syllabic structure. Syllables are claimed to be the most basic units in speech; every language has syllables, and children learn to produce them before they can say full words. Syllable structure can be of three types: simple (CV), moderate (CVC), and complex (with consonant clusters at edges), such as CCVCC or CCCVCC (where V = vowel and C = consonant). Words can be monosyllabic (one syllable), bisyllabic or disyllabic (two syllables), trisyllabic (three syllables), or polysyllabic (many syllables).
Topic-149: Stress as a Suprasegmental Feature
Stress is a suprasegmental feature applied to a whole syllable when it is made prominent by adding factors such as loudness, a rise in pitch, length of duration, and vowel quality (in contrast with other syllables). For example, in mi-ni-mi-ZA-tion, the second last (penultimate) syllable is prominent because greater energy is applied to it; it is louder, longer, and has different quality and pitch than the other syllables.
📌 Example: Compare IN-sult (noun) vs. in-SULT (verb). In IN-sult, the first syllable is stressed, marking it as a noun. In in-SULT, the second syllable is stressed, marking it as a verb. Similarly: be-LOW vs. BILL-ow.
Stress is an important feature in both phonetics and phonology. Despite extensive study, there are many areas of disagreement among experts. It is almost certainly true that in all languages, some syllables are stronger than others; these have the potential to be described as stressed. Stress plays an important role in conveying (and changing) meaning.
Topic-150: Types and Categories of Stress
One area of some agreement is the levels of stress. Some descriptions manage with just two levels (stressed and unstressed), while others use more. In English, taking the word in-di-ca-tor as an example, the first syllable is the most strongly primary stressed, the third syllable is the next most strongly secondary stressed, and the second and fourth syllables are tertiary stressed (or unstressed). This gives three levels: primary, secondary, and tertiary stress.
🔑 Definition — Primary Stress: The strongest level of stress on a syllable in a word; it can change meaning and grammatical category (e.g., RE-cord (noun) vs. re-CORD (verb)).
🔑 Definition — Secondary Stress: A weaker stress level, typically used for better pronunciation but not for changing meaning (e.g., in in-di-ca-tor, the third syllable).
Primary stress is very important in English because it can change meaning and the grammatical category of words (e.g., INsult (noun) vs. inSULT (verb)). In long words like minimization, one syllable receives primary stress while one or more syllables receive secondary stress. Secondary stress is just for better pronunciation and does not change meaning. Another division is made between lexical stress (phonemic in nature) and sentence level (emphatic) stress.
Topic-151: Lexical and Emphatic Stress
In terms of its linguistic function, stress is treated under two different headings: word (lexical) stress and sentence (emphatic) stress.
🔑 Definition — Lexical stress: Stress applied at the syllable level, where only one syllable in a word receives primary stress. It has the ability to change the meaning and grammatical category of a word, as in 'IMport' (noun) vs. 'imPORT' (verb).
🔑 Definition — Emphatic (Sentence) stress: Stress applied to one word (rather than a syllable) in a sentence, making that word more prominent than the rest. This type of stress plays a role in intonation patterns and rhythmic features of the language, showing specific emphasis on the stressed word (which may highlight some information in the typical context).
💡 Why this matters: Sentence stress allows speakers to highlight specific information in a context, changing the implied meaning of an entire sentence by emphasizing different words.
📌 Example: To perceive sentence level stress, read the following sentences with shifting stress on the words in bold and judge the shift in emphasis:
- Did YOU drive to Peshawar last weekend? (Emphasizing the person, not someone else)
- Did you DRIVE to Peshawar last weekend? (Emphasizing the mode of travel)
- Did you drive to PESHAWAR last weekend? (Emphasizing the destination)
- Did you drive to Peshawar LAST weekend? (Emphasizing the time)
⭐ Key Takeaways
The syllable is the most basic unit of speech, consisting of an optional onset, a mandatory nucleus (usually a vowel), and an optional coda. Its structure determines the phonotactic possibilities of a language. Stress is a suprasegmental feature that makes a syllable prominent through loudness, pitch, length, and vowel quality. Stress levels are categorized as primary, secondary, and tertiary, with primary stress being crucial for changing word meaning and grammatical category (e.g., INsult vs. inSULT). Finally, a critical distinction exists between lexical stress (at the word level, which changes meaning) and sentence/emphatic stress (at the sentence level, which highlights specific information and affects rhythm and intonation).
🧠 Quick Revision Questions
- What are the three possible parts of a syllable, and which one is mandatory?
- Give an example of how primary stress can change the meaning and grammatical category of a two-syllable English word.
- Name at least three phonetic factors that contribute to making a syllable stressed.
- What is the difference between lexical stress and emphatic (sentence) stress?
- According to the lecture, what are the three levels of stress, and which one is primarily responsible for changing word meaning?
📘 Lecture 30 — SUPRASEGMENTAL FEATURES-II
📖 Overview: This lecture explores the rhythmic classification of languages into stress-timed and syllable-timed categories, examining their defining characteristics, examples, and the phonetic principles behind each. It also introduces pitch as a crucial suprasegmental feature, linking frequency of vocal fold vibration to auditory sensation and meaning in language. Understanding these rhythmic patterns and pitch phenomena is essential for analyzing the prosodic structure of speech across languages.
🗂️ Topics Covered
The lecture covers six main topics: Stress Timed Languages (defining the category where stressed syllables recur at regular intervals and dominate rhythmic patterns), Explaining Stress Timed Languages (detailing the concept of isochronism and how stress divides syllables at word and sentence levels), Syllable Timed Languages (where all syllables tend to have equal duration and occur at regular time intervals), Explaining Syllable Timed Languages (discussing isosyllabism and the debate over the accuracy of this rhythmic classification), and Pitch as a Suprasegmental Feature (exploring pitch as an auditory sensation linked to fundamental frequency, its role in tonal languages and intonation, and distinguishing voiced from voiceless sounds).
📝 Lecture Summary
Topic-152: Stress Timed Languages
Languages of the world are broadly divided into two rhythmic categories: stress timed languages and syllable timed languages, based on their modes of timing. In stress timed languages, stress dominates the rhythmic feature, meaning these languages seem to be timed according to stressed patterns — the division among syllables is made on the basis of stressed and unstressed patterns. Examples include English and German. In such languages, stressed syllables occur at regular intervals, and their units of timing are perceived accordingly. Stress-timed rhythm is characterized by a tendency for stressed syllables to occur at equal intervals of time.
🔑 Definition — Stress timed languages: Languages where stressed syllables occur at regular intervals, and the division among syllables is made on the basis of stressed and unstressed patterns.
📌 Example: English and German are classic examples of stress-timed languages, where the time between stressed syllables remains roughly equal regardless of how many unstressed syllables fall between them.
💡 Why this matters: This rhythmic classification affects how learners perceive and produce speech rhythms in a foreign language, impacting intelligibility and naturalness.
Topic-153: Explaining Stress Timed Languages
This term characterizes the pronunciation of languages displaying a rhythmic pattern opposed to syllable-timed languages. In stress-timed languages, stressed syllables recur at regular intervals of time (stress-timing) regardless of the number of intervening unstressed syllables, as in English. This characteristic is sometimes referred to as isochronism or isochrony. However, this regularity occurs only under certain conditions, and the extent to which English exhibits this tendency compared to other Germanic languages remains unclear. In short, the division among syllables is made on the basis of stress and unstressed patterns. Stress is realized both at word and sentence levels, approximately changing rhythmic patterns, particularly at sentence level.
🔑 Definition — Isochronism/Isochrony: The property of stressed syllables recurring at regular intervals of time in stress-timed languages.
📌 Example: In the sentence "The doctor said he would come soon," the stressed syllables (doc, said, come, soon) occur at roughly equal time intervals, even though the number of unstressed syllables between them varies.
Topic-154: Syllable Timed Languages
In syllable timed languages, all syllables tend to have an equal time value (their length or duration), and the rhythm is said to be syllable-timed. Syllables tend to occur at regular intervals of time with fixed word stress. A classic example is Japanese, where all morae (phonological units) have approximately the same duration. This tendency is contrasted with stress-timing, where the time between stressed syllables is said to be equal irrespective of the number of unstressed syllables in between. Czech, Polish, Swahili, and Romance languages (e.g., Spanish and French) are often claimed to be syllable-timed. However, many phoneticians doubt whether any language is truly syllable-timed.
🔑 Definition — Syllable timed languages: Languages where all syllables tend to have equal duration, and syllables occur at regular intervals of time with fixed word stress.
📌 Example: In French, each syllable tends to take roughly the same amount of time to produce, so a word like "extraordinaire" (ex-tra-or-di-naire) would have each syllable lasting approximately the same duration.
Topic-155: Explaining Syllable Timed Languages
Syllable timed is a term used to characterize languages displaying a rhythmic pattern opposed to stress-timed languages. In these languages, syllables occur at regular intervals of time (e.g., in French). This characteristic is sometimes referred to as isosyllabism or isosyllabicity. However, very little work has been done on the accuracy or general applicability of such properties, and the usefulness of this typology has been questioned. It is important to remember that the division of languages into two categories (syllable vs. stress timed) is not based on strong concepts. Many phoneticians disagree with the basic idea of timing value, proposing instead three dimensions: fixed word stress (mainly in Romance languages), variable word stress (mainly in English and German), and fixed phrase stress (as exhibited by Japanese). They want to categorize languages on the basis of these three patterns.
🔑 Definition — Isosyllabism/Isosyllabicity: The property of syllables occurring at regular intervals of time in syllable-timed languages.
📌 Example: In Spanish (a syllable-timed language), each syllable in "Me gusta mucho" (Me / gus / ta / mu / cho) would have approximately equal duration, unlike English where stressed syllables dominate.
Topic-156: Pitch as a Suprasegmental Feature
As a suprasegmental feature, pitch is an auditory sensation — when we hear a regularly vibrating sound (like a note on a musical instrument or a vowel produced by the human voice), we hear a high pitch (when the rate of vibration is high) and a low pitch (when the rate of vibration is low). Some speech sounds are voiceless (e.g., /s/) and cannot give rise to a sensation of pitch, but voiced sounds can. The pitch sensation from a voiced sound corresponds closely to the frequency of vibration of the vocal folds. We usually refer to the vibration frequency as fundamental frequency to keep these two concepts distinct. In tonal languages, pitch is an essential component of word pronunciation, and a change of pitch may cause a change in meaning. In most languages (whether or not they are tone languages), pitch plays a central role in intonation. In simple words, pitch is the variation in the vibration of vocal folds.
🔑 Definition — Pitch: An auditory sensation corresponding to the frequency of vibration of the vocal folds, perceived as high or low.
🔑 Definition — Fundamental frequency: The actual physical measurement of vocal fold vibration frequency, distinct from the perceived pitch sensation.
📌 Example: In Mandarin Chinese (a tonal language), the syllable "ma" can mean "mother" (high level tone), "hemp" (rising tone), "horse" (falling-rising tone), or "scold" (falling tone), depending entirely on the pitch pattern used.
💡 Why this matters: Pitch distinguishes meaning in tonal languages and conveys attitude, emotion, and grammatical information through intonation in all languages, making it a critical feature for communication.
⭐ Key Takeaways
Students must remember that stress-timed languages (like English and German) have stressed syllables occurring at regular intervals regardless of intervening unstressed syllables (isochronism), while syllable-timed languages (like French, Spanish, Japanese) have all syllables with approximately equal duration (isosyllabism). The stress-timed vs. syllable-timed classification is debated among phoneticians, with some proposing three dimensions (fixed word stress, variable word stress, fixed phrase stress) instead of two. Pitch is a suprasegmental auditory sensation corresponding to vocal fold vibration frequency (fundamental frequency), and it plays a crucial role in both tonal languages (where pitch changes word meaning) and intonation in all languages. Voiceless sounds cannot produce pitch, while voiced sounds do, and pitch perception is closely tied to fundamental frequency.
🧠 Quick Revision Questions
- What distinguishes stress-timed languages from syllable-timed languages in terms of rhythmic patterns?
- Explain the concept of isochronism and provide an example from English.
- Why do many phoneticians question the accuracy of dividing languages into only stress-timed and syllable-timed categories?
- What is the relationship between pitch and fundamental frequency, and how do they differ?
- How does pitch function differently in tonal languages versus non-tonal languages?
📘 Lecture 31 — SUPRASEGMENTAL FEATURES-I II
📖 Overview: This lecture explores suprasegmental features, focusing on tone and intonation. It distinguishes between tonal and non-tonal languages, explains the nature and functions of intonation in communication, and discusses how these features convey grammatical structure, personal attitude, and discourse meaning.
🗂️ Topics Covered
The lecture begins by defining tone as a suprasegmental feature and examining tonal languages, using Mandarin Chinese as an example. It then defines intonation in both narrow and broad senses, describing various analytical approaches including contour and tone unit analyses. The major functions of intonation are detailed, including its grammatical, attitudinal, and discourse roles. Finally, the lecture explores intonation's function in conversational discourse, including turn-taking and the signaling of new versus old information.
📝 Lecture Summary
Topic-157: Tone and Tonal Languages
Tone (in phonetics and phonology) as a suprasegmental feature refers to an identifiable movement (variation) or level of pitch that is used in a linguistically contrastive way. In tone (tonal) languages, the linguistic function of tone is to change the meaning of a word. For example, in Mandarin Chinese, [́ma] said with a high pitch means ‘mother’ while [̀ma] said on a low rising tone means ‘hemp’. In other (non-tonal) languages, tone forms the central part of intonation, and the difference between, for example, a rising and a falling tone on a particular word may cause a different interpretation of the sentence in which it occurs. In the case of tone languages, it is usual to identify tones as being a property of individual syllables, whereas an intonational tone may be spread over many syllables. In the analysis of English intonation, tone refers to one of the pitch possibilities for the tonic (or nuclear) syllable. For further analysis, a set of four types of tone is usually used (fall, rise, fall–rise and rise–fall) though others are also suggested by various experts.
🔑 Definition — Tone: An identifiable movement or level of pitch used in a linguistically contrastive way. 📌 Example: In Mandarin Chinese, [́ma] (high pitch) = ‘mother’, [̀ma] (low rising tone) = ‘hemp’.
Topic-158: Intonation as a Suprasegmental Feature
In simple sense, intonation refers to the variations in the pitch of a speaker’s voice used to convey or alter meaning (at sentence level). In its broader and more popular sense, it is used to cover much the same field as prosody, where various features such as voice quality, tempo and loudness are also included. It is a term frequently used in the study of suprasegmental phonology, referring to the distinctive use of patterns of pitch, or melody and the study of intonation is sometimes called intonology. Experts have suggested several ways of analyzing intonation. In some approaches, the pitch patterns are described as contours and analyzed in terms of levels of pitch as pitch phonemes and morphemes while in others, the patterns are described as tone units or tone groups which are further analyzed as contrasts of nuclear tone, tonicity, etc. The three variables of pitch range, height and direction are generally distinguished. Some approaches, especially within pragmatics, operate with a much broader notion than that of the tone unit: intonational phrasing is a structured hierarchy of the intonational constituents in conversation. A formal category of intonational phrase is also sometimes recognized which is an utterance span dominated by boundary tones.
🔑 Definition — Intonation: Variations in the pitch of a speaker’s voice used to convey or alter meaning at the sentence level.
Topic-159: Functions of Intonations
Intonation as a suprasegmental feature performs several functions in a language. Its most important function is to act as a signal of grammatical structure (e.g., creating patterns to distinguish among grammatical categories), where it performs a role similar to punctuation (in written language). It may furnish far more contrasts (for conveying meaning). Intonation also gives an idea about the syntactic boundaries (sentence, clause and phrase level boundaries). It also provides the contrast between some grammatical structures (such as questions and statements). For example, the change in meaning illustrated by ‘Are you asking me or telling me?’ is regularly signaled by a contrast between rising and falling pitch. Note the role of intonation in sentences like ‘He’s going, isn’t he?’ (= I’m asking you) opposed to ‘He’s going, isn’t he!’ (= I’m telling you) (These examples are given by Peter Roach). The role of intonation in the communication is quite important as it also conveys personal attitude (e.g., sarcasm, puzzlement, anger, etc.). Finally, it can signal contrasts in pitch along with other prosodic and paralinguistic features. It can also bring variation in meaning and can prove an important signal of the social background of the speakers.
💡 Why this matters: Intonation is not just about melody; it is a critical grammatical and pragmatic tool that distinguishes questions from statements, signals speaker attitude, and even provides social information about the speaker.
📌 Example: ‘He’s going, isn’t he?’ (rising pitch = asking) vs. ‘He’s going, isn’t he!’ (falling pitch = telling).
Topic-160: Explaining Functions of Intonation
While discussing the functions of intonation, another approach is to concentrate on its role in conversational discourse. This involves such aspects as indicating whether the particular thing being said constitutes new information or old (in sentence, for example). It further creates the regulation of turn-taking in conversation, the establishment of dominance and the elicitation of co-operative responses as well. As with the signaling of attitudes, it seems that though analysts concentrate on pitch movements there are many other prosodic factors being used to create these effects. It is also important to note here that much less work has been done on the intonation of languages other than English. It seems that all languages have something that can be identified as intonation (there can be many differences between among languages). It is, therefore, a potential area to be studied and explored further. One reason for intonation not being well explored is because of the different descriptive frameworks used by different analysts for studying intonation cross-linguistically. Even it is claimed that tone languages also have intonation and that it is superimposed upon the tones themselves. Such a claim creates especially difficult problems of analysis.
⭐ Key Takeaways
A student must remember that tone uses pitch to distinguish word meanings, as in tonal languages like Mandarin Chinese, while intonation uses pitch variations at the sentence level to convey grammar and attitude. Intonation functions as a signal of grammatical structure (e.g., distinguishing questions from statements), can mark syntactic boundaries, and conveys personal attitudes like sarcasm or anger. In discourse, intonation helps regulate turn-taking and signals whether information is new or old. Finally, all languages have intonation, but it is a challenging area for cross-linguistic study due to different descriptive frameworks and the complex interaction of tone and intonation in tonal languages.
🧠 Quick Revision Questions
- What is the key difference between a tone language (like Mandarin Chinese) and a non-tonal language (like English) in how pitch is used linguistically?
- What are the four types of tones typically used in the analysis of English intonation?
- Name three distinct functions of intonation as discussed in the lecture.
- Provide the example from Peter Roach that demonstrates how intonation can change a tag question from a genuine question to a statement.
- Why is the study of intonation across different languages considered a challenging area of research?
📘 Lecture 32 — Linguistic Phonetics-I
📖 Overview: This lecture introduces linguistic phonetics as the approach embodied in the principles of the International Phonetic Association (IPA). It explains the distinction between the phonetics of the community and the phonetics of the individual, emphasizing the importance of shared phonetic knowledge. The lecture also provides a detailed overview of the IPA chart, its components, and how it represents humanly possible speech sounds.
🗂️ Topics Covered
This lecture covers five main topics: an introduction to linguistic phonetics as a framework for formal phonological theory; the distinction between the phonetics of the community and the phonetics of the individual; an explanation of why community phonetics is the primary focus; a description of the International Phonetic Alphabet (IPA) and its chart components; and an explanation of how the IPA chart reflects linguistic phonetics, including blank and shaded cells.
📝 Lecture Summary
Topic-161: Linguistic Phonetics
Linguistic phonetics is an approach embodied in the principles of the International Phonetic Association (IPA) and in a hierarchical phonetic descriptive framework that provides a basis for formal phonological theory. Speech, being a very complex phenomenon with multiple levels of organization, needs to be explored from different angles. Linguistic phonetics answers questions related to the possible ways of articulatory unified phonetics and phonology and from the perspective of cognitive phonetics, focusing on speech production and perception and how they shape languages as sound systems. The idea is mainly related to the overall ability of human beings to produce sounds (as a community and irrespective of their specific languages) and then the representation of their shared knowledge (as considered by the IPA in its charts) for formal phonetic and phonological theories.
Topic-162: Phonetics of the Community and of the Individual
Linguistic phonetic descriptions of speech sounds are, by and large, descriptions of the phonetics of the community, excluding the individual properties and considering the shared properties of the sound system of a language by its native speakers. The representations that experts write and use in the IPA, and analyze in a formal phonological theory, are intended to show the community’s shared knowledge of how to say the words of a language. This shared phonetic knowledge is perceptible to other speakers and to the phonetician. Experts are mainly concerned with the aggregate behavior of the linguistic group, capturing what community members accept as the correct pronunciation system. The focus of linguistic phonetic description is the phonetics of the community, while the phonetics of the individual is considered only for specific purposes when required.
🔑 Definition — Phonetics of the Community: The shared phonetic knowledge and pronunciation patterns accepted by native speakers of a language as correct, excluding individual variations. 🔑 Definition — Phonetics of the Individual: The phonetic knowledge and skills related to the individual's performance of language, including private knowledge and the role of memory and experience.
Topic-163: Explaining Phonetics of the Community and of the Individual
The major reason why the phonetics of the community is considered for phonetic descriptions is that individual speakers differ in interesting ways—two native speakers of a language will always speak with some variations. Describing the phonetics of the individual involves describing the phonetic knowledge and skills related to the performance of language. It is possible that certain aspects of the phonetics of the individual can be captured using IPA transcription, but others are not compatible with it (such as private knowledge, its performance, and the role of memory and experience). Secondly, the phonetics of the individual is usually not the focus of the linguist in speech elicitation, and it is difficult to describe even with spectrograms of the person’s speech. Although the phonetics of the individual is the focus of much of the explanatory power of phonetic theory, for general phonetic description we need to focus on the phonetics of the community.
💡 Why this matters: This distinction ensures that linguistic phonetic descriptions are generalizable and capture the systematic patterns of a language, rather than being overwhelmed by the infinite variations of individual speakers.
Topic-164: The IPA
While discussing the key elements of linguistic phonetic description, we need to consider the International Phonetic Alphabet (IPA). IPA is the set of symbols and diacritics that have been officially approved by the International Phonetic Association. The association publishes a chart comprising a number of separate charts. At the top inside the front cover, you will find the main consonant chart. Below it is a table showing the symbols for nonpulmonic consonants, and below that is the vowel chart. Inside the back cover is a list of diacritics and other symbols, and a set of symbols for suprasegmental features such as tone, intonation, stress, and length. The IPA chart does not try to cover all possible types of phonetic descriptions (e.g., all individual strategies for realizing linguistic phonological contrasts, or gradations in the degree of co-articulation between adjacent segments). Instead, it is limited to those possible sounds that can have linguistic significance in that they can change the meaning of a word in some languages. The description of IPA is based on the linguistic phonetics of the community.
🔑 Definition — International Phonetic Alphabet (IPA): The set of symbols and diacritics officially approved by the International Phonetic Association, used to represent the sounds of human speech.
Topic-165: Explaining the IPA
One evidence that the IPA chart is based on linguistic phonetics is the description of the blank cells on the chart (those neither shaded nor containing a symbol). These blank cells indicate combinations of categories that are humanly possible but have not been observed so far to be distinctive in any language (e.g., a voiceless retroflex lateral fricative is possible but has not been documented so far, so it is left blank). The shaded cells, on the other hand, exhibit sounds not possible at these places. Further, below the consonant chart is a set of symbols for consonants made with different airstream mechanisms (clicks, voiced implosives, and ejectives). All these descriptions reflect the potentialities of human speech sounds (as a linguistic community), not only showing the possible segments but also the suprasegmental features and points related to the possible airstream mechanisms and even the diacritics for various types of co-articulations and secondary articulatory gestures. The IPA chart is carefully documented by experts and is continuously revised and updated.
💡 Why this matters: The presence of blank and shaded cells demonstrates that the IPA is grounded in empirical observation—it documents what is humanly possible and what has been attested in languages, rather than being a purely theoretical construct.
⭐ Key Takeaways
The most critical point from this lecture is that linguistic phonetics focuses on the phonetics of the community—the shared, perceptible phonetic knowledge of a language's speakers—rather than on individual variations. This approach is embodied in the IPA chart, which documents and represents only those sounds with linguistic significance that can change meaning in some language. The IPA chart includes separate sections for consonants, nonpulmonic consonants, vowels, diacritics, and suprasegmental features like tone and stress. Importantly, blank cells on the chart represent humanly possible but unattested sounds, while shaded cells represent impossible combinations of articulatory categories. Students must remember that the IPA is continuously revised and updated by experts to reflect our growing understanding of human speech.
🧠 Quick Revision Questions
- What is the primary difference between the phonetics of the community and the phonetics of the individual?
- Why is the phonetics of the community preferred for general phonetic description over the phonetics of the individual?
- What do blank cells on the IPA chart represent, and what is an example provided in the lecture?
- What are the main components of the IPA chart as described in this lecture?
- What does it mean that the IPA is limited to sounds with "linguistic significance"?
📘 Lecture 33 — Linguistic Phonetics-II
📖 Overview: This lecture explores the feature hierarchy in linguistic phonetics, detailing how speech sounds are classified based on hierarchical features. It also critically examines the limitations of purely linguistic explanations and introduces the role of individual phonetic knowledge in explaining sound patterns like assimilation. Understanding the feature hierarchy is crucial for formally describing and classifying speech sounds.
🗂️ Topics Covered
The lecture begins by introducing the concept of feature hierarchy and how features are organized to define ever more specific phonetic properties. It then discusses the major regions of the vocal tract (Labial, Coronal, Dorsal, Radical, Glottal) and their subdivisions. Next, the explanation of feature hierarchy covers manner categories (Stop, Fricative, Approximant, Vowel) and their dependent features, including laryngeal and airstream features. Finally, the lecture addresses problems with linguistic explanations, using assimilation as an example to contrast community-level phonetics with individual-level phonetic knowledge.
📝 Lecture Summary
Topic-166: Feature Hierarchy
The feature hierarchy is an important concept in phonetics and phonology based on the properties and features of sounds. A feature may be tied to a particular articulatory maneuver or acoustic property. For example, the feature [bilabial] indicates that the segment is produced with both lips. Such features are listed in a hierarchy with nodes defining ever more specific phonetic properties. Sounds are first divided in terms of their supra-laryngeal and laryngeal characteristics, and their airstream mechanism. The supra-laryngeal characteristics can be further divided into those for place (of articulation), manner (of articulation), the possibility of nasality, and the possibility of being lateral.
Topic-167: Feature Hierarchy: Discussion
The first division in the feature hierarchy is made on the basis of the major regions of the vocal tract, giving five features: Labial, Coronal, Dorsal, Radical, and Glottal. The first three (Labial, Coronal, Dorsal) are related to tongue position. Radical is a cover term for [pharyngeal] and [epiglottal] articulations made with the root of the tongue. The feature Glottal is based on being [glottal], to cover various articulations such as [h]. Supra-Laryngeal features must allow for the dual nature of the actions of the larynx and include Glottal as a place of articulation. A sound may be articulated at more than one of these five regions. Within the five general regions, Coronal articulations can be split into three mutually exclusive possibilities: Laminal (blade of the tongue), Apical (tip of the tongue), and Sub-apical (under part of the blade of the tongue). Thus, the major regions may be subdivided into sub-regions.
Topic-168: Feature Hierarchy: Explanation
This grouping under feature hierarchy also reflects that the four features (Stop, Fricative, Approximant, and Vowel) depend on the degree of closure of the articulators. The manner category Stop has only one possible value ([stop]), but Fricative has two ([sibilant] and [nonsibilant]). Approximant and Vowel have five principal features: Height (with five possible values: [high], [mid-high], [mid], [mid-low], and [low]), Backness (with three values: [front], [center], and [back]), and two kinds of Rounding — Protrusion (with values [protruded] and [retracted]) and Compression (with values [compressed] and [separated]). The fourth feature for vowels and approximants is the Tongue Root, which has two possible values: [+ATR] and [-ATR]. Finally, the feature Rhotic has only one possible value: [rhotacized].
The Laryngeal possibilities involve mainly five features: [voiceless], [breathy voice], [modal voice], [creaky voice], and [closed] (forming a glottal stop). Airstream features have three values: Pulmonic, Velaric, and Glottalic.
💡 Why this matters: This hierarchical system provides a formal and precise framework for classifying all speech sounds, which is essential for phonological analysis and understanding sound patterns across languages.
Topic-169: Problems with Linguistic Explanations
The linguistic phonetic explanation does not always satisfy all topics of speech production, and we sometimes need to consider the phonetics of the individual. Current research focuses on speech motor control, the representation of speech in memory, and the interaction of speech perception and production and their roles in language change. Take the example of assimilation, where adjacent sounds come to share phonetic properties. If we restrict ourselves to the terminology of linguistic phonetics (the phonetics of the community), we are restricted to descriptions of sound patterns and not their explanation. Explanations restricted in this way often fall into the fallacy of reification (acting as if abstract things are concrete).
Topic-170: Explaining the Problem with Linguistic Explanations
The phonetics of the individual gives explanation to a number of complex issues. For assimilation, it happens because there is a tendency in pronunciation for adjacent sounds to share phonetic properties (at the individual level). This “explanation” can be stated as a formal constraint in Optimality Theory on sequences of sounds:
📐 Formula: AGREE(x): Adjacent output segments have the same value of the feature x. → In simple words, this means that assimilation exists when adjacent segments share features. Adjacent sounds influence sounds in the neighborhood, and the tendency to assimilate (a cross-linguistic generalization) exists because of a tendency to assimilate (reified as a specific “explanatory principle”). The private phonetic knowledge of the individual provides a more satisfying way to explain language sound patterns.
💡 Why this matters: This distinction highlights the limitations of purely abstract linguistic models and emphasizes the need to consider individual speaker performance and cognitive processes for a complete explanation of sound patterns.
⭐ Key Takeaways
The feature hierarchy organizes phonetic features hierarchically, starting with major vocal tract regions (Labial, Coronal, Dorsal, Radical, Glottal) and subdividing them into more specific properties. Manner features (Stop, Fricative, Approximant, Vowel) depend on the degree of articulatory closure and have specific possible values for height, backness, rounding, tongue root, and rhoticity. Laryngeal features include five voice qualities, and airstream features include pulmonic, velaric, and glottalic mechanisms. A key challenge is that linguistic phonetic explanations can fall into the fallacy of reification, describing patterns without explaining them, whereas the phonetics of the individual provides more satisfying explanations for phenomena like assimilation.
🧠 Quick Revision Questions
- What are the five major vocal tract regions in the feature hierarchy, and which one is a cover term for pharyngeal and epiglottal articulations?
- Into which three mutually exclusive possibilities can Coronal articulations be split?
- What are the five possible values for the feature Height in vowels and approximants?
- Name the five laryngeal features mentioned in the lecture.
- Using assimilation as an example, explain the difference between a linguistic phonetic explanation and an explanation based on the phonetics of the individual.
📘 Lecture 34 — Linguistic Phonetics-I II
📖 Overview: This lecture explores the phonetic dimensions of individual speech production, focusing on how articulatory movements are controlled and how speech is stored in memory. It addresses the balance between speaker and listener needs, explaining the phonetic forces that shape sound patterns and historical changes in pronunciation.
🗂️ Topics Covered
The lecture examines controlling articulatory movements through muscular complexity, memory for speech and the lack of phonetic invariance, exemplar theory's explanation of speech memory, the balance between phonetic forces like ease of articulation and perceptual separation, and how languages maintain this balance through assimilation and contrast maintenance.
📝 Lecture Summary
Topic-171: Controlling Articulatory Movements
This section explores the speech motor control underlying individual phonetic production. For example, producing a voiceless bilabial [p] involves an array of muscular complexity including dozens of muscles in the chest, abdomen, larynx, tongue, throat, and face — all contracted with varying degrees of tension in specific sequence and duration. For lip closure, two main muscles (depressor labii inferior and incisivus inferior) are activated to create 'too much' and 'enough tension'. Simultaneously, jaw muscles trade with lip muscles for closing and opening. This structure specifies an overall task "close the lips" at the top node, with subtasks like "raise the lower lip" and "lower the upper lip" coordinated to accomplish the overall goal. Additionally, for voiceless bilabial [p], the glottis needs to be wide apart for free air passage. Thus, exploiting the air passage, keeping it voiceless at larynx, and creating closure by lips and jaws with many subtasks are achieved through controlling articulatory movements, enabling understanding of individual variation in speech production.
💡 Why this matters: This shows that even a simple sound like [p] involves complex motor coordination, explaining why individual speakers produce sounds differently.
Topic-172: Memory for Speech
Speech is diverse and complex, particularly regarding the phonetics of the individual. Different speakers of the same language will have somewhat different productions depending on vocal tract physiology, their own habits of speech motor coordination, and their memory of speech. Sociolinguistic features also influence production — we are exposed to speech styles ranging from careful pronunciations in public speaking to casual styles typical between friends. This leads to lack of phonetic invariants (or variability of invariability). This 'lack of phonetic invariance' poses an important problem for phonetic theory as we try to reconcile that shared phonetic knowledge can be described using IPA symbols and phonological features with the fact that individual phonetic forms span a very great range. This problem has great practical significance for language engineers trying to get computers to produce and recognize speech. The solution is facilitated by 'the phonetic implementation approach' that focuses on how experiences are encoded in memory. According to this view, words are stored in speech memory in their most basic phonetic form and used when needed.
🔑 Definition — Lack of phonetic invariance: The phenomenon where the same linguistic unit (e.g., a phoneme) is produced with highly variable acoustic and articulatory characteristics across different contexts, speakers, and speaking styles.
Topic-173: Explaining Memory for Speech
The role of memory for speech under exemplar theory suggests that many instances of each word are stored in memory, and their phonetic variability is memorized rather than computed. The main postulates include:
- Language universal features: Broad phonetic classes (e.g., aspirated vs. unaspirated) derive from physiological constraints on speaking or hearing, but their detailed phonetic definitions are arbitrary — a matter of community norms.
- Speaking styles: No one style is basic (from which others are derived), because all are stored in memory. Bilingual speakers store two systems.
- Generalization and productivity: Exemplar theory says generalization is possible within productivity. However, productivity — the hallmark of linguistic knowledge in the phonetic implementation approach — is the least developed aspect of exemplar theory.
- Sound change: Sound change is phonetically gradual and operates across the whole lexicon. It is a gradual shift as new instances keep on adding.
🔑 Definition — Exemplar theory: A theory of memory for speech which posits that all experienced instances (exemplars) of words are stored in memory, and phonetic categories emerge from the similarity structure of these stored tokens.
Topic-174: The Balance Between Phonetic Forces
To explain the sound patterns of a language, the views of both speaker and listener are considered. Both prefer to use the least possible articulatory effort (except when producing very clear speech), resulting in numerous assimilations, with some segments left out and others reduced to minimum. Thus a speaker uses language with ease of articulation (e.g., co-articulation and secondary articulation). This tendency toward maximum ease of articulation leads to change in pronunciation of words. For example, in co-articulations, a change in the place of the nasal and following stop occurred in words like improper and impossible before these words came into English through Norman French. In these words, the [n] from the prefix in- (as in intolerable and indecent) changed to [m], reflected even in spelling. These are historical cases of assimilation, where one or more segments are affected by adjacent segments for economy of articulation. Articulation with least possible effort enables us to speak by keeping the balance between phonetic forces, ultimately leading to changes in pronunciation.
💡 Why this matters: Historical sound changes like in- becoming im- before labials show how ease of articulation shapes language over time, even affecting spelling.
Topic-175: Explaining the Balance Between Phonetic Forces: Explanation
When producing sounds with maximum ease of articulation, only similar sounds are affected. The focus of speakers is always on maintaining a sufficient perceptual distance between sounds that occur in a contrasting set (e.g., vowels in stressed-monosyllabic words beat, bit, bet, and bat). This principle of perceptual separation does not usually result in one sound affecting an adjacent sound (as in maximum ease of articulation). Instead, perceptual separation affects the set of sounds that potentially can occur at a given position in a word — such as the vowel position in stressed monosyllables — so that perceptual separation is maximized. The principle of 'maximum perceptual separation' also accounts for some differences between languages. These examples illustrate how languages maintain a balance between the requirements of the speaker (pressure for easier articulations) and those of the listener (sufficient perceptual contrast between sounds that affect meaning).
🔑 Definition — Maximum perceptual separation: The principle that sounds in a contrasting set (e.g., vowels in minimal pairs) tend to be maintained at sufficient perceptual distance to preserve distinctiveness and avoid confusion in meaning.
⭐ Key Takeaways
The lecture emphasizes that speech production involves complex motor control where multiple muscles coordinate hierarchically to achieve articulatory goals like lip closure for [p]. Memory for speech, explained through exemplar theory, stores many instances of words with their phonetic variability rather than computing them, which accounts for lack of phonetic invariance across speakers and styles. Languages balance two opposing phonetic forces: the speaker's drive for ease of articulation (producing assimilations like [n] becoming [m] before labials) and the listener's need for perceptual separation between contrasting sounds (e.g., maintaining distinct vowels in beat, bit, bet, bat). This balance explains both synchronic variation and historical sound change, with phonetic implementation approach and exemplar theory providing frameworks for understanding how individual phonetic forms relate to shared linguistic knowledge.
🧠 Quick Revision Questions
- What muscular actions are involved in producing the lip closure for [p], and how do subtasks coordinate?
- What is the "lack of phonetic invariance" and why is it a problem for phonetic theory and language engineering?
- According to exemplar theory, how are words stored in memory and how does this account for phonetic variability?
- How does the principle of ease of articulation explain the historical change from [n] to [m] in words like improper?
- What is the principle of maximum perceptual separation and how does it balance speaker and listener needs in contrasting sets like beat, bit, bet, bat?
📘 Lecture 35 — PHONOTACTICS AND SYLLABIC TEMPLATES
📖 Overview: This lecture explores phonotactics and syllabic templates, focusing on how languages organize sounds into syllables and the permissible sequences of phonemes. It explains syllabification, the universal CV syllabic pattern, and how languages classify syllable structures as simple, moderate, or complex based on consonant clusters. Understanding these concepts is crucial for analyzing phonological well-formedness and cross-linguistic variation.
🗂️ Topics Covered
The lecture covers three main topics: Syllabic Templates and Syllabification, where the syllable as a unit of pronunciation is defined and the universal CV pattern is introduced; Explaining Syllabification, which discusses internal syllable structure, consonant clusters, and the classification of languages into simple, moderate, and complex syllabic patterns; and Phonotactics, which examines the study of permissible sound sequences in languages, including sequential constraints, phonotactic rules, and examples from English phonotactics.
📝 Lecture Summary
Topic-176: Syllabic Templates and Syllabification
A syllable is the unit of pronunciation typically larger than a single sound and smaller than a word. Words may be divided into parts, as in ne-ver-the-less, and dictionaries indicate syllabic divisions for hyphenation. Syllabification refers to the division of a word into syllables, while resyllabification is a reanalysis that alters syllable boundary locations. A word with one syllable is monosyllabic; with more than one, it is polysyllabic. The CV pattern (one consonant at the onset followed by a vowel as its peak) is found in all languages of the world — it is the universal syllable pattern (Max Onset Principle). Some languages only allow CV templates, e.g., Honolulu (CVCVCVCV) and Waikiki (CVCVCV). This pattern is also common in nicknames across languages: kami, nana, baba, papa, mani, rani. Children first acquire the CV pattern of their mother tongue during L1 acquisition.
Topic-177: Explaining Syllabification
Languages differ in their internal syllable structure, and syllabic patterns vary across languages and families. Consonant sequences are called clusters (e.g., CC — two consonants or CCC — three consonants). Most phonotactic analyses are based on syllable structures and syllabic templates. On the basis of consonant clusters at edges (onset and coda), three types of syllabic patterns are identified:
🔑 Definition — Simple syllabic pattern: CV (e.g., to, be) 🔑 Definition — Moderate syllabic pattern: CVC(G)(N) — where G stands for Glide and N stands for Nasal (specific consonants only) 🔑 Definition — Complex syllabic pattern: CCVCC to CCCVCC — involving bipartite (CC) and tripartite (CCC) clusters
Topic-178: Phonotactics
Phonotactics is the study of phonemes and their order found in syllables — the study of sound sequences in a language. Languages do not allow all phonemes to appear in any order. For example, an English speaker knows that /streŋθs/ makes the English word strengths and that /bleidʒ/ would be acceptable as blage (though it does not exist), but /lvm/ could not be part of an English word.
Phonotactic analyses reveal interesting findings: Why do words like bump, lump, hump, rump, mump(s), clump associate with large blunt shapes? Why do words ending with a plosive and a syllabic /l/ (e.g., muddle, fumble, straddle, cuddle, fiddle, buckle, struggle, wriggle) relate to clumsy, awkward, or difficult action? Why can't English syllables begin with /pw/, /bw/, /tl/, /dl/ when /pl/, /bl/, /tw/, /dw/ are acceptable? All such discussion is called phonotactics of the language.
179. Explaining Phonotactics
Phonotactics refers to the order (sequential arrangements or tactic behavior) of segments (sounds or phonological units) that occur in a language. It shows what counts as a phonologically well-formed structure of a word. The allowed and restricted sound patterns of language are found through phonotactics. For example, in English, consonant sequences like /fs/ and /spm/ do not occur initially in words, and there are many other restrictions on possible consonant+vowel combinations. By analyzing data, sequential constraints of a language can be stated as phonotactic rules.
According to Generative phonotactics, no phonological principles can refer to morphological structure; phonological patterns sensitive to morphology (e.g., affixation) are represented only in the morphological component of grammar, not in phonology.
Examples of English phonotactic patterns:
| Pattern | Example |
|---|---|
| One phoneme (V) | I, oh, owe |
| Two phoneme (CV) | to, be, see |
| Three phoneme (CVC) | cat, dog, run |
| Four phoneme (CCVC) | stick, click, brick |
| Five phoneme (CCVCC) | brisk, treats, speaks |
| Six phoneme (CCCVCC) | streets, strand, strips |
| Seven phoneme (CCCVCCC) | strengths |
Other possible patterns: CCV (try), CCCVC (stroke), CCCV (straw), VCC (eggs), CVCC (risk), CVCCC (risks).
⭐ Key Takeaways
The most critical points from this lecture are: (1) The CV pattern is the universal syllable template found in all languages, and children acquire it first; (2) Syllable structures are classified as simple (CV), moderate (CVC(G)(N)), or complex (CCVCC to CCCVCC) based on consonant clusters at onset and coda; (3) Phonotactics studies permissible sound sequences in a language, revealing both allowed and restricted patterns; (4) Languages impose sequential constraints — certain consonant clusters are allowed initially (e.g., /pl/, /tw/) while others are not (e.g., /pw/, /tl/) in English; (5) Generative phonotactics holds that phonological principles cannot refer to morphological structure — patterns sensitive to morphology belong to the morphological component, not phonology.
🧠 Quick Revision Questions
- What is the universal syllable pattern found in all languages, and why is it significant?
- How are languages classified based on their syllabic patterns, and what do the terms "simple," "moderate," and "complex" refer to?
- Define phonotactics and provide two examples of permissible and two examples of impermissible English syllable-initial consonant clusters.
- According to Generative phonotactics, where should morphological sensitive phonological patterns be represented?
- Give one example each of a five-phoneme, six-phoneme, and seven-phoneme English word pattern, and explain what each pattern indicates about consonant clusters.
📘 Lecture 36 — Using PRAAT-I
📖 Overview: This lecture introduces the PRAAT software, a free and powerful tool for phonetic analysis, synthesis, and manipulation of speech. It is designed to equip students with fundamental skills for recording, displaying, segmenting, labeling, and exporting speech data, which are essential for acoustic phonetics research.
🗂️ Topics Covered
The lecture begins with an introduction to PRAAT, its manual, and basic exploration of the software. It then guides students through the practical steps of recording and displaying speech, segmenting and labeling sound files using textgrids, and exporting visual displays of waveforms and spectrograms to a Word document for reporting purposes.
📝 Lecture Summary
Topic-180: Introduction to PRAAT
PRAAT is a freeware program created by Paul Boersma and David Weenink at the Institute of Phonetic Sciences, University of Amsterdam. It can be freely downloaded from http://www.praat.org or http://www.fon.hum.uva.nl/praat/. The software allows you to analyze, synthesize, and manipulate speech, and create high-quality pictures for articles and theses. Active discussions and blogs are available on the 'yahoo Praat group', and an introductory tutorial is on the homepage. 💡 Why this matters: PRAAT is the standard tool in modern phonetics and phonology research, replacing many expensive alternatives.
🔑 Definition — PRAAT: A computer program used for analyzing, synthesizing, and manipulating speech, and for creating high-quality visual displays.
Topic-181: Introduction to PRAAT Manual
The manual used in this course was developed by Sonya Bird and Qian Wang at the University of Victoria, Canada. It contains carefully designed labs with worksheets. This course will conduct important experiments from the first five labs, enabling students to continue independently. PRAAT requires only modest computer experience, and queries can be addressed to the developers or the 'yahoo Praat group'. It is essential for learning acoustic phonetics and future research work.
Topic-182: Exploring PRAAT
Start by downloading PRAAT from www.praat.org based on your PC or Mac specifications. Initially, explore the two main windows: the Object window and the Picture window. To begin, double-click on the icon to open PRAAT and get to know its layout, including buttons on the Object window, main menu, analysis and synthesis tools, and the Picture window options.
Topic-183: Recording and Displaying
This is the first practical experiment. To record a new mono sound:
- Go to NEW > Record mono-sound (with a sampling rate of 44100 Hz).
- Ensure the volume bar fluctuates; if not, you're not recording or speaking loudly enough.
- Watch out for clipping: if the recording level goes into the red on the volume bar, you'll get a clipped signal, which is bad for speech analysis.
- Give the recording a name in the "Save to list" box and click "Save to list."
To display the recording:
- In the Objects window, highlight the sound file and click Edit (on the "Analysis and synthesis tools" panel).
- To extract a portion for study, select the desired part, then go to File > Extract Selection. The extracted selection appears as 'sound untitled' in the Objects window. Close the Edit window.
🔑 Definition — Clipping: A distortion that occurs when the recording level is too high, causing the signal to be cut off at the peaks; it is detrimental to speech analysis.
Topic-184: Segmenting and Labeling
Segmenting and labeling is crucial for acoustic analysis, allowing you to place segmental symbols and add textgrids to spectrograms. Follow these steps:
-
Create a textgrid: Highlight the sound file in the Objects window, then go to Annotate > To TextGrid. Create two tiers (e.g., 'word segment') and enter them in the 'All tier names' cell.
-
Open sound file and textgrid together: Hold down Ctrl and click on both files to highlight them. Click Edit to see the waveform (top), spectrogram (middle), and textgrid (bottom).
-
Segment the file: Place the cursor at the beginning of a target name on the spectrogram/waveform to show a boundary line. Click in the little circle at the top of the word tier to create a boundary. To remove a boundary, highlight it and go to Boundary > Remove or hit Alt+backspace.
-
Label the intervals: Select/highlight the target interval by clicking between two boundaries (the interval turns yellow). To input or change text, edit in the Textbox above the spectrogram. Give each interval a name. For phonetic symbols, go to the TextGrid editing window and select Help > Phonetic symbols.
🔑 Definition — Textgrid: An annotation file in PRAAT that contains tiers with boundaries and intervals for labeling segments of a sound file.
Topic-185: Exporting Visual Display to Word File
Exporting the visual display to Word is part of the experiment's write-up. To export your labeled waveform and spectrogram:
- Maximize the Edit window by clicking the square at the top right.
- Hit the PrtScr (Print Screen) button on your keyboard.
- Open a Word document, provide an introduction as required, and hit Ctrl+V to paste the PRAAT image.
- Give the figure a number and title below the image.
⭐ Key Takeaways
The critical points from this lecture are: PRAAT is a free and essential tool for acoustic phonetics used to analyze, synthesize, and manipulate speech. The first practical step is recording a mono sound at a 44100 Hz sampling rate while avoiding clipping. Segmenting and labeling require creating a textgrid with tiers, opening it alongside the sound file, and placing boundaries between intervals on the waveform/spectrogram. Finally, exporting the visual display to Word is done by maximizing the Edit window, using the Print Screen key, and pasting the image. Mastering these basics enables independent future work in phonetic research.
🧠 Quick Revision Questions
- What is PRAAT, and who created it?
- What is the recommended sampling rate for recording mono sound in PRAAT, and why is clipping harmful?
- Describe the steps to create a textgrid and open it together with a sound file for editing.
- How do you create and remove a boundary on a textgrid tier?
- What is the simple method to export a labeled waveform and spectrogram from PRAAT into a Word document?
📘 Lecture 37 — Using Praat-II
📖 Overview: This lecture focuses on the practical application of Praat for acoustic analysis of speech sounds, particularly sonorous sounds like vowels. It explains the Source-Filter Model of speech production and provides step-by-step instructions for measuring fundamental frequency, harmonics, and formants using Praat software. Understanding these measurements is crucial for analyzing the acoustic components of speech.
🗂️ Topics Covered
This lecture covers the Source-Filter Model of speech, followed by three measurement techniques: measuring fundamental frequency (F0/pitch) using three methods, measuring harmonics using narrow-band spectrograms, measuring formants using wide-band spectrograms and three techniques, and finally exploring the relationship between harmonics and formants. The lecture emphasizes the distinction between harmonics (source/laryngeal) and formants (filter/vocal tract).
📝 Lecture Summary
Topic-186: The Source-Filter Model of Speech
The Source-Filter theory is crucial for understanding the basic components of speech sounds and the nature of acoustic signals (the physics of speech sounds). It is particularly important for acoustic analysis of vowels and vowel-like sonorous sounds. The major goal is to understand and explore basic acoustic components of sonorous speech sounds such as fundamental frequency, harmonics, and formants. These components create the acoustic signals associated with speech, and understanding them is crucial to understanding what we actually hear.
Topic-187: Measuring the Fundamental Frequency
Fundamental frequency (F0 or Pitch) is an important component of source-filter theory. It can be taken from the middle of the sound (e.g., vowel) using three techniques: displaying the Pitch track (automatically), measuring from Pitch track manually, and by looking at the waveform by highlighting one cycle.
Method 1: Displaying the pitch track and allowing Praat to measure pitch automatically:
- Display the pitch track: Pitch > Show pitch.
- Place your cursor in the middle – a stable portion of the vowel.
- Go to Pitch > Get pitch – a box will appear with the pitch value in it.
Method 2: Displaying the pitch track and measuring pitch manually:
- Display the pitch track - Pitch > Show pitch.
- Click on the blue pitch track in the middle of the vowel.
- A red horizontal bar should appear with the pitch value (in dark blue) on the right side of the window.
Method 3: By looking at the waveform (top of the display):
- Zoom into a small piece of the waveform in the middle of the vowel and measure the period by highlighting one complete cycle and noting the time associated with it.
Topic-188: Measuring the Harmonics
Harmonics are the multiple integers of the fundamental frequency which are basically the result of vocal fold vibration (complex wave). We need the narrow band spectrogram for measuring the H (which we can set by fixing the spectrum setting at 0.025). By starting measuring the frequency of the first three harmonics, we go to the H10 (H1, H2, H3 – H10). When vocal folds vibrate, the result is a complex wave, consisting of the fundamental frequency plus other higher frequencies called harmonics.
To see harmonics, we need a narrow-band spectrogram, which is more precise along the frequency domain than the default wide-band spectrogram.
Procedure:
- Display a narrow-band spectrogram: Spectrum > Spectrogram settings.
- Change the window length to 0.025s (the default window length is 0.005s for wide-band spectrogram).
- Notice the grey horizontal bands corresponding to harmonics.
- For each vowel, measure the frequencies of the first 3 harmonics (H1-H3) and the 10th harmonic (H10).
- Click on the center of each harmonic in the center of each vowel.
- A red horizontal bar should appear with the frequency value on the left side of the window in red.
🔑 Definition — Harmonics: Multiple integers of fundamental frequency that result from vocal fold vibration (complex wave).
📐 Formula: Harmonics = n × F0 (where n = integer, F0 = fundamental frequency) → Harmonics are whole-number multiples of the fundamental frequency.
Topic-189: Measuring the Formants
Formants are the overtone resonances. Acoustically, to plot vowels on a chart, F1-F2 are very important. Narrow band spectrograms are required for measuring harmonics whereas we need the wide bands for measuring formants (important characteristics of sonorant speech sounds – vowels). On spectrogram, formants are thick bands (darkness corresponds to loudness; the darkest harmonics are the most amplified). These amplified harmonics form the formants characteristic of sonorant speech sounds.
Method 1: Displaying the formants (red dots on the spectrogram) automatically:
- Display the formant track: Formant > Show formants.
- Place cursor in the middle, stable portion of the vowel.
- Go to Formant > Formant listing: a box appears with time point and first four formants.
Method 2: Displaying the formants and measuring the frequency manually:
- Display the pitch track: Formant > Show formants.
- Place cursor in the center of each formant, in the middle of the vowel.
- A red horizontal bar appears with the frequency value on the left side of the window in red.
Method 3: Measuring the frequency without displaying Praat formants (if Praat's formant tracking goes wonky):
- Remove Praat's formant tracking: Formant > Show formants (unclick).
- Place cursor in the center of each formant, in the middle of the vowel.
- A red horizontal bar appears with the frequency value on the left (in red).
🔑 Definition — Formants: Overtone resonances; amplified harmonics that are characteristic of sonorant speech sounds.
💡 Why this matters: F1-F2 are essential for plotting vowels on a vowel chart, which is fundamental for acoustic phonetic analysis.
Topic-190: Relationship Between Harmonics and Formants
Having taken measurements for both formants and harmonics, we compare them to explore the relationship between the two. Captured in the Source-Filter Model of speech, from comparison of the two values, it is clear that harmonic numbers are different for one type of sounds but the formants are the same. The relationship is: harmonics are related to the laryngeal activity (source) and formants are the output of the vocal tract (filter).
The figure illustrates: Source (vocal folds produce complex wave with harmonics) → Filter (vocal tract shape amplifies certain harmonics to create formants) → output speech sound.
⭐ Key Takeaways
Fundamental frequency (F0) can be measured using three methods: automatic pitch display, manual pitch tracking, or waveform cycle measurement. Harmonics are integer multiples of F0 resulting from vocal fold vibration and are measured using narrow-band spectrograms (window length 0.025s). Formants are amplified harmonics shaped by the vocal tract and are measured using wide-band spectrograms; the first two formants (F1, F2) are essential for vowel identification. The Source-Filter Model explains that harmonics come from the laryngeal source while formants are produced by the vocal tract filter. A key observation is that harmonic numbers vary across sounds but formant frequencies remain consistent for the same vowel.
🧠 Quick Revision Questions
- What are the three methods for measuring fundamental frequency in Praat?
- What spectrogram setting is required for measuring harmonics, and why?
- What is the relationship between harmonics and formants according to the Source-Filter Model?
- Why are F1 and F2 measurements particularly important in acoustic phonetics?
- How can you measure formants if Praat's automatic formant tracking is unreliable?
📘 Lecture 38 — Using Praat-III
📖 Overview: This lecture guides students through the practical acoustic analysis of vowels using Praat software. It covers measuring intrinsic pitch (F0) and the first two formants (F1 and F2) for eight American English vowels, culminating in plotting these values on a chart in Excel. This hands-on session is critical for connecting acoustic measurements to articulatory properties of vowels.
🗂️ Topics Covered
The lecture comprises four main topics: first, exploring vowel properties by recording eight specific American English vowels. Second, measuring intrinsic pitch (F0) for each vowel using Praat's tools. Third, analyzing the spectral make-up by measuring the first two formant values (F1 and F2). Fourth, plotting the formant values on a chart in Excel, including reversing axes to create the traditional vowel space.
📝 Lecture Summary
Topic-191: Vowel Properties
To explore the acoustics of vowels, you need to record the eight vowels from American English: heed, hid, head, had, hod, hawed, hood, and who’d. After recording, you will measure three things: intrinsic pitch, spectral make up (formants), and plot them in an Excel sheet. First, look at your vowels in the Edit window to ensure you can clearly see the vowel formants; if not, review previous labs.
Topic-192: Intrinsic Pitch
Measure the pitch (F0) for each of the eight recorded vowels using any of the three methods discussed in Topic-187. If values seem strange, use another method for confirmation—measuring the waveform is the safest way. To confirm pitch via a spectral slice:
- Select a 70-80ms portion of the vowel.
- Go to Spectrum > View spectral slice.
- Click on the first (big) peak = H1 = F0 (ignore small spikes, which may be noise). Note the frequency of this peak at the top of the vertical bar.
Use the confirmed pitch values to plot the pitch of each vowel on your Excel sheet. Label the y-axis with a scale that spreads out measurements. Create a cluster chart, export it to Word, and give it a figure number and title.
Topic-193: Spectral Make-up (Formant Values)
To determine the spectral make-up of the vowels, measure the first two formants (F1 and F2) for each of the eight vowels. Use either automatic formant tracking or manual measurement (trust your judgment over Praat's). Use the default wide-band spectrogram for measuring F1 and F2. Also, answer these questions by revisiting chapters on acoustics:
- a. Why do formants (F1 and F2) differ across vowels?
- b. What does F1 correspond to in terms of articulation?
- c. What does F2 correspond to in terms of articulation?
💡 Why this matters: These questions directly link acoustic measurements (formants) to articulatory features (tongue height and frontness), forming the basis of vowel classification.
Topic-194: Plotting Vowels on Chart
Using the F1 and F2 measurements, plot the values on an Excel chart. Organize the data in three columns: vowels in the first column, difference between F2 and F1 in the second column, and F1 in the third column. Use the Scatter chart type. To make the chart correspond to the traditional vowel space, reverse the values for both formants on both axes (Y and X), so that zero for F1 and F2 is at the right corner. After plotting, note that F1 is inversely related to vowel height, and the difference between F2 and F1 is related to vowel frontness. Export the chart to Word and give it a number and title.
⭐ Key Takeaways
The critical skills from this lecture are recording and analyzing the eight American English vowels in Praat, measuring both intrinsic pitch (F0) and the first two formants (F1 and F2). You must be able to confirm pitch measurements using spectral slices and use the wide-band spectrogram for manual formant tracking. The final, crucial step is plotting F1 against the F2-F1 difference in Excel, reversing both axes to produce the standard vowel space chart. This process directly demonstrates that F1 corresponds to vowel height (open/close) and F2-F1 difference to vowel frontness (front/back).
🧠 Quick Revision Questions
- What are the eight American English vowels (as words) you must record for this lecture?
- Describe the "safest" method for measuring pitch (F0) if Praat's automatic values seem strange.
- How can you use a spectral slice to confirm your pitch (F0) measurement?
- In the Excel chart for formants, which two values do you plot on the axes, and which axis is reversed to create the traditional vowel space?
- What articulatory property does F1 correspond to, and what property does the difference between F2 and F1 correspond to?
📘 Lecture 39 — Using PRAAT - IV
📖 Overview: This lecture explores the acoustic analysis of sonorants (nasals, glides, and liquids) using Praat software. It focuses on measuring formants and identifying distinctive acoustic patterns that differentiate these sounds from vowels, which is essential for understanding acoustic phonetics.
🗂️ Topics Covered
The lecture covers sonorants and their formants, nasal formants and their distinctive waveforms including anti-formants and formant transitions, and glide formants and their comparison with vowels. It guides students through hands-on recording and analysis of sequences like /ama/, /ana/, /aŋa/, /wi/, and /ju/ to explore systematic acoustic patterns.
📝 Lecture Summary
Topic-195: Sonorants and Their Formants
Sonorants are vowel-like sounds, specifically nasals and glides. They are called sonorants because they have formants (the acoustic correlates of resonance). However, they differ from vowels because they generally have lower amplitude, which makes them behave like consonants. To analyze these sounds, record the sequences: /ama/ - /ana/ - /aŋa/ - /wi/ - /ju/.
🔑 Definition — Sonorants: Vowel-like consonants (nasals and glides) that have formants but lower amplitude than vowels.
💡 Why this matters: Distinguishing sonorants from vowels is crucial for identifying consonant types in spectrograms.
Having recorded these sequences, explore the features by measuring F1, F2, and F3 for each sound. Compare these measurements with those of vowels. Look for systematic patterns in the formant values before proceeding to the next topic.
Topic-196: Nasal Formants
Formants for nasal sounds are essential for acoustic analysis. Measure the first three formants (F1, F2, and F3) of nasals from the recorded files, using the previously learned method of measuring formants. Remember that nasals have very distinctive waveforms that differ from vowels. They have distinctive forms of anti-formants, which are bands of frequencies that are damped, and formant transitions.
🔑 Definition — Anti-formants: Bands of frequencies that are damped or reduced in amplitude, characteristic of nasal sounds.
When measurements are complete, answer these questions:
- Are there any systematic patterns across nasals?
- Is there one formant with a similar frequency for all places of articulation?
- Is there one formant that has much higher amplitude than the others across nasals?
- Do you see any overall differences between the nasals on the one hand and [a] on the other?
📌 Example: Comparing /m/, /n/, and /ŋ/ in the sequence /ama/-/ana/-/aŋa/ — the F1 value might be consistently around 250-300 Hz for all nasals (answer to Q2), while F2 varies with place of articulation (e.g., lower F2 for /m/ due to labial closure).
Topic-197: Glide and Their Formants
Glides, such as /w/ and /j/, are also sonorants that are vowel-like because they have formants. From the recorded files, take the first three formants (F1, F2, and F3) from the middle of the sounds (the midpoint of [w] and [j]). Analyze how similar the formant structure of glides is to that of vowels and nasals. Draw lines to indicate F1, F2, and F3 on the spectrogram and compare them with vowels.
🔑 Definition — Glides: Sonorant consonants (e.g., /w/ and /j/) that have formant structures similar to vowels but function as consonants in syllable structure.
After analysis, answer these questions:
- Which vowel does each glide resemble? (e.g., /w/ resembles [u] and /j/ resembles [i])
- Are there any differences between the contours, i.e., between the transitions out of the two glides?
- What do you think may cause these differences (in terms of articulation)?
📌 Example: For /wi/ — the F1 and F2 of [w] will be low (like [u], F1300Hz, F2700Hz), while the [i] part will show a rapid transition to high F2 (~2300Hz). For /ju/ — [j] will show low F1 (~300Hz) but high F2 (~2300Hz, like [i]), transitioning to lower F2 for [u].
⭐ Key Takeaways
This lecture emphasizes that sonorants (nasals and glides) are vowel-like consonants with formants but lower amplitude than vowels, making them acoustically distinct. Nasals are identified by their anti-formants and distinctive waveforms, while glides show formant structures that resemble specific vowels (/w/ like [u], /j/ like [i]). Systematic patterns across nasals include a consistent formant frequency for all places of articulation, and glides exhibit formant transitions that differ based on their articulatory origins. Understanding these acoustic features is critical for distinguishing sonorants from vowels in spectrographic analysis.
🧠 Quick Revision Questions
- What makes sonorants different from vowels in terms of amplitude?
- What are anti-formants, and why are they important for identifying nasal sounds?
- Which vowel does the glide /w/ resemble, and which does /j/ resemble, based on formant structure?
- When measuring formants of nasals, what is the first step in finding systematic patterns?
- What causes the differences in formant contours between the two glides /w/ and /j/?
📘 Lecture 40 — Using PRAAT-V
📖 Overview: This lecture provides a practical, hands-on guide to analyzing acoustic features of obstruents (stops, fricatives, and affricates) using PRAAT software. It focuses specifically on how to measure and interpret voicing correlates in stops, including the voice bar, Voice Onset Time (VOT) , and preceding vowel duration, enabling students to distinguish voiced, voiceless, and aspirated stops acoustically.
🗂️ Topics Covered
This lecture covers the acoustic analysis of stop consonants using PRAAT. It begins with an examination of stop voicing on spectrograph, focusing on the voice bar and the duration of preceding vowels. It then moves to measuring Voice Onset Time (VOT) , with detailed steps for calculating and comparing negative, zero, and positive VOT across voiced, voiceless, and aspirated stops like /apa/, /aba/, /ata/, /ada/, /apha/, and /atha/. The lecture concludes with guided questions for self-assessment.
📝 Lecture Summary
Topic-198: Stop Voicing on Spectrograph
There are three important acoustic correlates of voicing in stops: the voice bar, VOT (Voice Onset Time), and the duration of the preceding vowel. To analyze these, record tokens like /apa/, /aba/, /ata/, /ada/, /apha/, and /atha/. For each stop in the file, take three measurements: first, see the voicing or the voice bar by exploring features of the stop. Second, explore features related to the place of articulation (e.g., any bilabial feature for /p/ or /b/ in comparison with non-bilabial sounds). Third, check the duration of the preceding vowels. Note down the presence of voicing. The voice bar appears as a low-frequency energy band during the closure period of a voiced stop. If the voice bar disappears, it may be due to a lack of glottal vibration or a very short closure duration.
🔑 Definition — Voice Bar: A low-frequency energy band visible on a spectrogram during the closure of a voiced stop, indicating ongoing vocal fold vibration. 📐 No formula; this is a visual measurement technique. 📌 Example: When analyzing the word /aba/ in PRAAT, zoom into the spectrogram during the /b/ closure. Look for a dark horizontal line near the bottom of the spectrogram (around 0–200 Hz). If it is present throughout the closure, the stop is fully voiced; if it disappears mid-closure, the voicing may have ceased.
✅ Guided Questions:
- Do you see the voice bar? If the voice bar is present at all, does it last through the duration of the closure? Why do you think it might go away? (Possible reason: subglottal pressure equalizes, stopping vocal fold vibration.)
- Can you point out the features related to the place of articulation? Comment on them (e.g., bilabial stops show a diffuse burst spectrum; alveolar stops show a more compact burst).
- How about the duration of the preceding vowel? Is it affected by the voicing? (Vowels are typically longer before voiced stops than before voiceless stops.)
💡 Why this matters: Distinguishing voiced from voiceless stops in a spectrogram is foundational for acoustic phonetic analysis and helps in understanding phonological contrasts.
Topic-199: Measuring Voice Onset Time (VOT)
Another aspect of acoustic correlate of stops is VOT, which is a characteristic of voiced, voiceless, and aspirated stop sounds. There are very easy steps to calculate VOT. Record /apa/, /aba/, /ata/, /ada/, /apha/, and /atha/. Zoom in through your stop sounds so that you can analyze the patterns and find the difference among the three types of VOT (negative, zero, and positive). Measure the VOT of each stop and compare voiced/voiceless counterparts (p/b, t/d, k/g). Similarly, zoom in so that you can clearly see the stop closure followed by the beginning of the vowel. You can measure the time between the end of the stop closure (the beginning of the release burst) and the onset of voicing in the following vowel (the onset of regular pitch pulses in the waveform). This measurement is Voice Onset Time or VOT.
🔑 Definition — Voice Onset Time (VOT): The time interval between the release of a stop closure (the burst) and the onset of vocal fold vibration (voicing) for the following vowel. 📐 Formula: VOT = Time(onset of voicing) − Time(release burst). Measured in milliseconds (ms). 📌 Example: For a voiceless unaspirated stop like /p/ in /apa/, the VOT is near zero (voicing starts almost immediately after the burst). For a voiced stop like /b/ in /aba/, VOT is negative (voicing begins during the closure, before the burst). For an aspirated stop like /ph/ in /apha/, VOT is positive and long (there is a delay of 50–100 ms between the burst and voicing onset).
✅ Guided Questions:
- How does VOT differ in voiced and voiceless stops and what articulatory explanation can you come up with for this? Answer: Voiced stops have negative or short positive VOT; voiceless stops have longer positive VOT. Articulation: In voiced stops, vocal folds are already adducted and vibrating; in voiceless stops, they are abducted, needing time to adduct and begin vibrating.
- How does the duration of the preceding vowel differ depending on the voicing of the following consonant? Answer: Vowels are longer before voiced stops than before voiceless stops (a phenomenon called pre-fortis clipping).
💡 Why this matters: VOT is a critical acoustic cue that helps listeners distinguish between phonemic categories like /b/ vs. /p/ in English, and its measurement is central to acoustic phonetics research.
⭐ Key Takeaways
The most critical concepts from this lecture for the exam are the three acoustic correlates of stop voicing: the voice bar (a low-frequency energy band during closure for voiced stops), Voice Onset Time (VOT) (the time between stop release and voicing onset, which can be negative, zero, or positive), and the duration of the preceding vowel (longer before voiced stops). Students must be able to measure VOT in PRAAT by zooming in on the waveform to identify the release burst and the onset of pitch pulses, and then interpret these values to classify stops as voiced, voiceless, or aspirated. The lecture emphasizes practical data collection and analysis using recordings of minimal pairs like /apa/ vs. /aba/.
🧠 Quick Revision Questions
- What are the three acoustic correlates of voicing in stops mentioned in this lecture?
- How do you identify a voice bar on a spectrogram, and what does its presence indicate?
- Define Voice Onset Time (VOT) and state the formula for measuring it from a waveform.
- Compare the typical VOT patterns for a voiced stop /b/, a voiceless stop /p/, and an aspirated stop /ph/.
- How does the duration of a preceding vowel change depending on whether the following stop consonant is voiced or voiceless?
📘 Lecture 41 — Further Areas of Study in P&P
📖 Overview: This lecture outlines potential and current research areas in Phonetics and Phonology (P&P), particularly within the Pakistani context. It emphasizes the importance of P&P research as part of English Language Teaching (ELT) and introduces other advanced topics like distinctive features, experimental phonetics, and the study of language varieties.
🗂️ Topics Covered
This lecture covers P&P research as part of ELT, focusing on issues faced by Pakistani learners and the documentation of Pakistani English. It also discusses current trends in P&P research on regional languages, the study of distinctive features using binary analysis, the field of experimental phonetics with its latest technological tools, and the study of language varieties through comparisons of accents and dialectology.
📝 Lecture Summary
Topic-200: P&P Research as the Part of ELT
Phonetics and phonology is a very potential area for research in the Pakistani context. In applied phonology, many studies can explore issues faced by Pakistani learners of English, such as pronunciation difficulties. Researchers can also document the phonological features of Pakistani English to get this variety recognized. Other problematic areas include segmental and suprasegmental features like stress placement, intonation, and syllabification. Contrastive analysis between English and regional languages (Urdu, Punjabi, Sindhi, Balochi, Pashto) is a rich research area, as is the study of consonant clusters and interlanguage phonology from a second language acquisition perspective. Further studies may focus on the development of corpora for Pakistani English and the application of IPA resources in ELT.
Topic-201: Current Trends in P&P Research
Current trends in P&P research extend beyond ELT to the documentation and study of regional languages of Pakistan. Researchers can get their work on IPA illustrations of these languages published in international journals like the IPA Journal of Cambridge University. The Himalaya Hindu Kush (HKH) region, one of the world's richest regions linguistically and culturally, is a very potential area for areal and typological linguistics. When working on these languages, one may also apply for funding from international organizations like those for endangered languages and UNESCO.
Topic-202: Distinctive Features
The study of distinctive features is a key area for phonetic studies. This phonological analysis describes a phoneme as a combination of different features in a binary (+/-) order, e.g., /d/ as [+alveolar, +stop, +voiced, +oral, +central]. The feature analysis must include three principles: contrastive function (how it is different), descriptive function (what it is), and classificatory function (based on broader classes of sounds). Features can also be studied as part of language universals and their role as language-specific subsets.
🔑 Definition — Distinctive Features: The smallest discrete units that make up a phoneme, defined in binary terms (e.g., +voiced, -voiced) to show contrasts. 📐 Formula: Binary Feature Analysis: /d/ = [+alveolar, +stop, +voiced, +oral, +central] 📌 Example: The phoneme /p/ would be described with features like [-voiced, +bilabial, +stop], distinguishing it from /b/ which is [+voiced].
Topic-203: Experimental Phonetics
Experimental phonetics involves the study of sounds using the latest experimental techniques and computer software under carefully designed lab experiments. This goes beyond simple acoustics by working in sophisticated phonetic labs to discover hidden aspects of human speech, such as how speech is produced and processed. The latest trends include studying brain functions in speech production and processing (using equipment like x-ray techniques), speech errors, neurolinguistics, and topics related to developments through computers for speech analysis and synthesis.
💡 Why this matters: Experimental phonetics connects linguistic theory with neuroscience and technology, offering a way to empirically verify models of speech production and perception.
Topic-204: The Study of Variety
The study of varieties of English (and other major languages) is a potential research area in P&P. This includes comparisons and contrasts among accents at phonetic and phonological levels, including segmental and suprasegmental features. English dialectology has been explored with a focus on geographic differences in the recognition of various forms of English (Englishes). Well-known data-gathering techniques from sociolinguistics, such as the variations paradigm, are used in the field. Field workers develop expertise for large-scale studies related to language varieties.
⭐ Key Takeaways
This lecture emphasizes the vast research potential in phonetics and phonology, particularly in the Pakistani context. The most critical areas include applied phonology for ELT, focusing on Pakistani English and learner difficulties; current trends like documenting regional languages for international journals; the theoretical framework of distinctive features using binary analysis; the advanced field of experimental phonetics using technology to study speech production and brain function; and the study of varieties through accent comparison and sociolinguistic methods.
🧠 Quick Revision Questions
- What are three specific pronunciation issues faced by Pakistani learners of English that can be studied in applied phonology?
- What is the significance of the Himalaya Hindu Kush (HKH) region for P&P research?
- What are the three principles for feature analysis in phonology?
- What are two latest trends or techniques used in experimental phonetics?
- What are the linguistic differences that the study of variety (dialectology) focuses on?
📘 Lecture 42 — The Pedagogy of Phonetics and Phonology
📖 Overview: This lecture explores the pedagogical implications of phonetics and phonology (P&P) for English Language Teaching (ELT), focusing on how teachers can effectively integrate P&P into their classroom practice. It provides practical guidance on developing relevant teaching materials, conducting classroom action research, and staying updated with the latest methodologies, specifically within the Pakistani context.
🗂️ Topics Covered
The lecture covers five main topics: the relationship between ELT and P&P, with a focus on integrating P&P into ELT and conducting phonological contrastive analysis; developing relevant material for teaching P&P, including using online resources and creating original content; conducting classroom research for ELT through action research to solve pedagogical problems; making research accessible to teachers to encourage continuous professional development; and facilitating action research to investigate specific issues like speech errors and learner performance.
📝 Lecture Summary
Topic-205: The Relationship Between ELT and P&P
This topic establishes that Phonetics and Phonology (P&P) is an integral part of English Language Teaching (ELT). The teaching of P&P must be integrated into ELT teaching. To achieve this, teachers are expected to improve their own English pronunciation skills and sensitize their students to the subject. They can use self-initiated procedures to carry out a phonological contrastive analysis (CA) — comparing, for example, the segmental and suprasegmental features of their students' mother tongues and English — to enhance teaching skills and complete research. Teachers are also encouraged to participate in phonology-based ELT activities from online sources like the TESOL Home Page, the English Language Teaching Reforms (ELTR) projects of the Higher Education Commission (HEC) of Pakistan, and activities sponsored by British Council Pakistan. Students should also be part of these platforms through social groups and online learning.
💡 Why this matters: This topic directly connects the theoretical study of phonetics and phonology to practical classroom application, showing teachers how to make it relevant and effective for their students.
Topic-206: Developing Relevant Material
Developing relevant material is a crucial task for aspiring teachers of English. Teachers should remember the specific needs of ELT activities in their own context and explore already developed material from online sources like the British Council. However, they must also be able to develop their own material tailored to their students' needs. For example, teachers can develop material for pronunciation teaching. This can include material related to IPA transcription of audio-based listening activities by involving students in using phonetic dictionaries in the classroom. Other effective resources include movies and documentaries from channels like BBC, CNN, and National Geographic. Finally, real-life material for listening and writing interaction from everyday language can also be very effective. The focus of material development should always be to enhance the students' proficiency level.
Topic-207: Conducting Classroom Research for ELT
Teachers are expected to engage in action research, a key aspect of ELT with significant pedagogical implications. In this context, action research means focusing on issues faced by ELT instructors and seeking solutions through research. English language learners in Pakistan face many problems, providing numerous opportunities to plan research inside the classroom. These can include problem-solving, material development, and using course-books like novels and dramas. Classroom action research is always based on problem-solving (e.g., teaching-related problems) and is mostly conducted by teachers themselves. As they work in the field, they are well aware of the problems faced by teachers and learners and are expected to plan authentic research on these issues and their solutions.
Topic-208: Making Research Accessible to Teachers
Good teachers are expected to be active researchers, constantly updating themselves on the latest teaching methodologies and research worldwide. A key pedagogical challenge is for teachers to stay updated by exploring pedagogical and technological challenges for ELT experts, both in their own contexts and internationally. For instance, the aspects of Task Based Learning and Teaching (TBLT) , considered a golden method for Second Language Acquisition (SLA) , could be effective in the Pakistani context if explored by ELT practitioners. Teachers are agents of change who must read research studies, conduct their own research, and explore their issues and solutions. A good way to do this is to regularly read teachers' digests and journals and participate in online discussions by teaching associations.
Topic-209: Facilitating Action Research
Teachers are expected to facilitate action research, which is the most rewarding and productive activity for their own profession. For example, the phonetics of phonological speech errors can be explored and shared by teachers investigating their own practices, leading to a very positive discussion in academic circles of research into ELT and SLA. Similarly, topics like learners' performance and development (e.g., "what do good speakers do?") can yield useful results for the teaching community. Teachers and student teachers are required to facilitate action research on reading and listening issues, English reading strategies in primary schools and their effectiveness, impact on pronunciation, and many more topics. Research in phonetic theory and description with phonological, typological, and broader implications can also be included in phonetics and phonology-specific action research.
⭐ Key Takeaways
The pedagogy of phonetics and phonology is inseparable from ELT, requiring teachers to actively integrate P&P into their practice. Teachers must be prepared to develop their own teaching materials, using authentic resources like movies and phonetic dictionaries, to meet their students' specific needs. A core responsibility of a language teacher is to become a practitioner of classroom-based action research, identifying and solving authentic pedagogical problems. To remain effective, teachers must actively seek out and engage with the latest research and teaching methodologies through professional networks and publications. The ultimate goal is to facilitate action research within their own classrooms to investigate learner performance, speech errors, and other critical areas that can improve the teaching and learning of English pronunciation.
🧠 Quick Revision Questions
- What are the key characteristics of the action research that ELT teachers are expected to conduct?
- Name three specific types of resources, as mentioned in the lecture, that can be used to develop effective pronunciation teaching materials.
- According to the lecture, how can teachers make research accessible to themselves for continuous professional development?
- What is the purpose of a teacher carrying out a phonological contrastive analysis (CA) in the context of teaching phonetics and phonology?
- The lecture describes facilitating action research on a specific phonetic topic. What is that topic, and what is its potential benefit for the academic community?