Acoustic Phonetics: Praat, Spectrograms & VOT

Learning Goal: Analyze vocalic and consonantal acoustic properties by reading spectrograms and measuring formants, pitch, and voice onset time (VOT) using Praat software.

  • Prerequisites: Basic familiarity with phonology/phonetics notation (IPA) is helpful but not strictly required. No prior physics or acoustic programming background is assumed.
  • Estimated Total Study Time: 17 Hours

Module 1: Foundations of Sound and Phonetics

In this module, you will learn how the human vocal tract acts as a biological instrument to produce speech. You will study articulatory mechanisms—voicing, place of articulation, and manner of articulation—and connect them to the fundamental physics of sound waves, laying the groundwork for reading these features on a spectrogram.

Recommended Videos

  • Why this video: This lecture establishes the high-level roadmap of phonetic sciences. It clearly divides phonetics into three distinct, interconnected branches: articulatory (how we make sounds), acoustic (the physical sound waves), and auditory (how we hear). This structural understanding is essential before you dive into the software-driven analysis of sound files.
  • Knowledge Checkpoint:
    • Explain the difference between articulatory, acoustic, and auditory phonetics.
    • Identify the anatomical structures involved in speech generation, including the vocal folds and oral/nasal cavities.
    • Define how sound propagates through the air as waves of compression and rarefaction.
  • Why this video: Before looking at waveforms in Praat, you must understand the "source" of those waveforms. This video breaks down the vocal tract anatomy in high detail, distinguishing between active articulators (like the tongue and lips) and passive articulators (like the alveolar ridge and hard palate). This physical model explains why certain acoustic signatures appear as they do.
  • Knowledge Checkpoint:
    • Distinguish between active and passive articulators and provide three examples of each.
    • Trace the path of air from the lungs through the glottis and out through the oral and nasal cavities.
    • Describe the acoustic consequences of modifying the shape of the vocal tract.
  • Why this video: This video introduces the three-part classification system for consonants: Voicing, Place of Articulation, and Manner of Articulation. It serves as a visual bridge, allowing you to connect physical muscle movements with the resulting sounds.
  • Knowledge Checkpoint:
    • Explain what physical action dictates whether a consonant is voiced or voiceless.
    • Identify the physical locations of bilabial, alveolar, and velar articulations.
    • Contrast stops, fricatives, and approximants based on how they restrict airflow.
  • Why this video: This video covers the primary articulatory axes of vowels: tongue height, backness, and lip rounding. It demonstrates why vowels are modeled as a continuous multi-dimensional space, which maps directly to the formant frequencies (F1 and F2) you will soon measure in Praat.
  • Knowledge Checkpoint:
    • Draw or describe the vowel trapezoid and locate the extreme point vowels: [i], [u], [a].
    • Explain how tongue height and backness change the interior volume of the oral and pharyngeal cavities.
    • Define how lip rounding alters the acoustic length of the vocal tract.

Module 2: Introduction to Spectrograms and Praat Software

This module transitions you from the anatomy of speech to digital acoustic analysis. You will download and navigate the Praat user interface, learn how to load and record audio, and understand the core differences between a waveform (time/amplitude) and a spectrogram (time/frequency/amplitude).

Recommended Videos

  • Why this video: This quick, practical tutorial introduces Praat's unique dual-window layout. It walks through the "Objects" window (where data is managed) and the "Picture" window (used for export/plotting). Knowing how to open and interact with these basic windows avoids common frustrations during your first software launch.
  • Knowledge Checkpoint:
    • Locate the "Objects" list and the "Picture" window when Praat opens.
    • Successfully load an external .wav audio file into the Praat Objects window.
    • Explain the purpose of the "Sound Editor" window and how to open it.
  • Why this video: This beginner-friendly guide walks through five fundamental operations in Praat, including uploading files and utilizing the live recording feature. It shows you how to capture your own speech for real-time analysis.
  • Knowledge Checkpoint:
    • Record a short phrase of your own voice using Praat's internal "Record mono Sound" function.
    • Identify and modify visual parameters inside the viewer, such as zooming in/out on specific sections.
    • Save analyzed audio files in .wav format to preserve acoustic integrity.
  • Why this video: This video bridges raw signals and visual analysis. It explains how a complex, raw sound wave is decomposed into its individual sine wave components using Fourier analysis, and how those components are mapped onto a three-dimensional spectrogram (Time on X, Frequency on Y, Amplitude as Darkness/Color).
  • Knowledge Checkpoint:
    • Differentiate between a waveform representation and a spectrogram representation of the same sound.
    • Identify what physical property is represented by the darkness or intensity of markings on a spectrogram.
    • Describe how a spectrum (slice of time) differs from a continuous spectrogram.

Module 3: Analyzing Vowels: Formants and Pitch

In this module, you will learn to read and measure vowel sounds. By using Praat’s visual markers, you will identify the fundamental frequency (F0F_0, perceived as pitch) and measure the first two vocal tract resonances, known as formants (F1F_1 and F2F_2). These measurements will allow you to plot and analyze a phonetic "vowel space."

High Vowels (Low F1) [i] (High F2) [u] (Low F2) \ / \ / \ / \ / \ / \ / [a] Low Vowels (High F1)

Recommended Videos

  • Why this video: This video introduces the "Source-Filter Theory." It teaches how the vocal cords generate a sound source (rich in harmonics) and how the physical shape of the oral/pharyngeal cavities acts as a filter, highlighting specific frequencies (formants) to create distinct vowel qualities.
  • Knowledge Checkpoint:
    • Define "formants" and explain why they are independent of the pitch (F0F_0) of your voice.
    • Explain how changing tongue height physically changes the size of the pharyngeal cavity, thus altering F1F_1.
    • Explain how tongue advancement (fronting vs. backing) alters oral cavity volume, thus altering F2F_2.
  • Why this video: This step-by-step tutorial shows you exactly where to click to find and measure formants in Praat. It introduces the red formant tracker dots, explains how to extract numerical values, and outlines the standard parameters for male vs. female speakers.
  • Knowledge Checkpoint:
    • Turn on the formant listing feature in the Praat editor window (Show Formants).
    • Measure the exact steady-state mid-point of F1F_1 and F2F_2 for three different vowel segments.
    • Adjust the maximum formant settings (e.g., 5000 Hz for adult males, 5500 Hz for adult females) to resolve tracker tracking errors.
  • Why this video: This advanced video demonstrates how to run prosodic evaluations in Praat. You will learn to isolate the blue pitch line (fundamental frequency, F0F_0) from vowel formants, adjust pitch floor/ceiling settings, and align these measurements with TextGrid segment annotations.
  • Knowledge Checkpoint:
    • Extract continuous pitch contour (F0F_0) data from a spoken sentence.
    • Correct pitch-tracking errors (such as octave jumps) by adjusting the pitch analysis range settings.
    • Create a basic TextGrid file with boundaries aligned to vowel segment boundaries.

Module 4: Analyzing Consonants: Acoustic Signatures

Consonants do not have the same clear, sustained bands of energy that vowels do. Instead, they present complex acoustic signatures: bursts, turbulence, silent intervals, and sudden shifts in energy. This module teaches you to identify stop gaps, fricative noise, nasal murmurs, and liquid transitions on a spectrogram.

Recommended Videos

  • Why this video: This in-depth video breaks down the distinct acoustic behaviors of various consonant classes. It demonstrates what silent intervals (stop closures), noise bands (frication), and transitional contours look like on a spectrogram, making it an excellent resource for visual identification.
  • Knowledge Checkpoint:
    • Identify a stop consonant's silent gap (closure) followed by its high-amplitude release burst.
    • Distinguish between the high-frequency, dark energy of a sibilant fricative (like [s]) and the low-frequency energy of a non-sibilant fricative (like [f]).
    • Describe the acoustic characteristics of an affricate, showing how it combines properties of stops and fricatives.
  • Why this video: This video uses clean, comparative graphics to teach spectrogram reading. It explains how to determine voicing status by looking for a voicing bar at the bottom of the spectrogram, and how to spot place-of-articulation clues in neighboring vowel transitions.
  • Knowledge Checkpoint:
    • Locate the "voicing bar" (low-frequency periodic energy) that characterizes voiced consonants.
    • Explain how a voiced stop gap differs visually from an unvoiced stop gap.
    • Describe how F2F_2 transitions in adjacent vowels provide visual clues about whether a consonant is bilabial, alveolar, or velar.
  • Why this video: This video focuses on the acoustic features of sonorants (nasals and approximants). It explains how nasalization introduces "antiformants" (damping of sound energy) and a distinctive low-frequency nasal murmur. It also covers how glides show dynamic formant movements.
  • Knowledge Checkpoint:
    • Identify a nasal murmur and explain why nasals appear noticeably lighter/fainter in upper frequencies than adjacent vowels.
    • Explain what "antiformants" (or zeroes) are and why they occur in nasal consonants.
    • Trace the formant transitions of approximants/glides ([j], [w]) and describe how they resemble rapid, non-stationary vowels.

Module 5: Measuring Voice Onset Time (VOT)

Voice Onset Time (VOT) is a key acoustic metric used to analyze phonological contrasts (such as voiced vs. voiceless stops). It measures the duration between the release of a stop consonant's burst and the onset of vocal fold vibration. This module covers how to measure this interval in Praat.

Unvoiced Aspirated Stop (e.g., [pʰ]): Waveform: |----- Stop Gap -----| BURST |==== Aspiration (VOT) ====|=== VOWEL === |<------ Measure VOT ----->|

Voiced Stop (e.g., [b]): Waveform: |~~~~ Voiced Gap ~~~~| BURST |=================== VOWEL ================ |<->| (Very short VOT, or negative VOT if pre-voiced)

Recommended Videos

  • Why this video: This tutorial directly addresses the core measurement goal of this curriculum. It provides a step-by-step walk-through in Praat, showing you how to locate the release burst, zoom in to the vowel onset, select the interval, and read the time duration in milliseconds.
  • Knowledge Checkpoint:
    • Zoom in to isolate a single stop-vowel sequence in the Sound Editor window.
    • Identify the exact visual start point of a stop release (the transient burst spike) on both the waveform and spectrogram.
    • Locate the first periodic glottal pulse of the following vowel to establish the end point of the VOT.
    • Calculate the final VOT value in milliseconds (or decimal seconds) from the Praat selection window.
  • Why this video: This video provides the theoretical context for why VOT matters in speech perception. It explains categorical perception and how the human auditory system uses a sharp VOT boundary (typically around 20-40 milliseconds) to divide a continuous acoustic signal into distinct phonemic categories like /b/ vs. /p/.
  • Knowledge Checkpoint:
    • Explain the concept of categorical perception in relation to VOT boundaries.
    • Identify typical VOT ranges for English voiced/unaspirated stops vs. voiceless/aspirated stops.
    • Define "pre-voicing" (negative VOT) and explain what it looks like on a spectrogram.
  • Why this video: This brief video demonstrates helpful interface shortcuts in Praat. It shows how to use TextGrids and the "SEL" (Select) button to quickly zoom into targets, streamlining your workflow when measuring large datasets.
  • Knowledge Checkpoint:
    • Use a TextGrid boundary to label a stop consonant.
    • Apply the "SEL" keyboard shortcut or button to zoom in on a highlighted interval instantly.
    • Verify that your highlighted selection spans from the first transient burst line to the first zero-crossing of the voiced periodic wave.

Step-by-Step Guide: Measuring VOT in Praat

Because VOT measurement requires high precision, follow these manual steps when analyzing your files:

  1. Locate the Stop: Look for a silent interval (stop gap) on the waveform, followed by a vertical spike on the spectrogram (the release burst).
  2. Find the Burst (Start Point): Click precisely at the start of this vertical burst spike. This is your start boundary (TstartT_{\text{start}}).
  3. Find the Voicing Onset (End Point): Look to the right of the burst. Find the first periodic wave cycle on the waveform that matches the vertical, repeating pitch striations on the spectrogram. Click at the point where this wave crosses the horizontal zero-amplitude line (zero-crossing). This is your end boundary (TendT_{\text{end}}).
  4. Read the Value: The duration of this highlighted selection (TendTstartT_{\text{end}} - T_{\text{start}}) is your VOT.
    • Aspirated Stops (e.g., [p], [t], [k]): Clear, noisy aspiration gap between burst and voicing. VOT is positive, typically >40 ms> 40 \text{ ms}.
    • Unaspirated Stops (e.g., [b], [d], [g] in English): Voicing begins almost immediately after the burst. VOT is small and positive, typically 020 ms0\text{--}20 \text{ ms}.
    • Pre-voiced Stops: Vocal fold vibration starts before the burst during the closure. This is negative VOT, represented by a voicing bar present throughout the stop gap.

Course Map

This flowchart shows the recommended learning path and module dependencies:


Key People Index

  • Paul Boersma & David Weenink: The phonetician-programmers at the University of Amsterdam who created Praat in 1992. Their work made speech analysis accessible worldwide, replacing expensive hardware setups with free software.
  • Gunnar Fant (1919–2009): A pioneer in acoustic phonetics who formulated the Source-Filter Theory of Speech Production in 1960. His research mathematically modeled how the vocal tract filters the source sound wave generated by the larynx.
  • Lisker & Abramson: Linguists who introduced Voice Onset Time (VOT) in 1964. They proved that VOT serves as a primary acoustic feature for distinguishing voiced and voiceless stops across different world languages.

Final Self-Assessment

Test your understanding of the course material by completing this checklist:

  • Draw a diagram of the vocal tract and label the active and passive articulators.
  • Explain the relationship between a complex sound wave's physical properties (frequency, amplitude) and how they appear on a spectrogram.
  • Successfully record, open, and navigate a mono audio file inside the Praat Sound Editor.
  • Correctly measure the fundamental frequency (F0F_0) of a vowel sound in Praat and adjust the pitch range to fix tracking errors.
  • Measure the first two formants (F1F_1 and F2F_2) of the vowels [i], [u], and [a] and explain how these measurements relate to tongue position.
  • Identify stop consonants, fricatives, nasals, and glides on a spectrogram using only visual cues.
  • Explain how antiformants and nasal murmurs are generated anatomically, and how they appear on a spectrogram.
  • Measure the Voice Onset Time (VOT) of voiceless aspirated vs. voiced unaspirated stops in Praat with millisecond precision.
  • Define negative VOT (pre-voicing) and describe how to identify it on a spectrogram.
  • Explain how categorical perception boundaries relate to how listeners categorize speech sounds using VOT cues.
Explore Further

Related Linguistics Roadmaps

View All