Speech Analysis: Sound Waves, Spectrograms & Formants

Added:

Waveform Basics
Spectrum Views
Early Spectrograph
Modern Spectrogram
Spectrogram Features
Formant Cavities
Vowel Chart
Consonant Spectra
Consonant Details
Final Summary

Waveform Basics

0:03
Playing Section
  • 1

    Explains waveform view showing amplitude over time.

  • 2

    Identifies silence and friction noise segments in speech.

Basic physics of sound waves, including the concepts of frequency (pitch), amplitude (loudness), and harmonics.
Fundamentals of articulatory phonetics, specifically how the vocal tract, tongue, and lips shape different speech sounds.
The core distinction between vowels (voiced, unobstructed airflow) and consonants (obstructed or restricted airflow) in human speech.
An introductory understanding of time-domain versus frequency-domain representations of signals.
Hands-on acoustic analysis using specialized software like Praat to measure fundamental frequency (F0) and formants (F1, F2) in real speech samples.
The role of acoustic features in Automatic Speech Recognition (ASR) and digital signal processing, such as Mel-Frequency Cepstral Coefficients (MFCCs).
Clinical applications of speech analysis in speech-language pathology to diagnose and monitor articulation, voice, and resonance disorders.
Forensic phonetics and speaker identification, using spectrograms and formant tracking for voice biometrics and legal investigations.
112.5K views1.6Klikes19:01@oer-vlcOriginal Release: 2013-09-06

Acoustic phonetics uses three main visualization tools—waveform views showing amplitude over time, frequency spectra displaying frequency-amplitude relationships at single time points, and spectrograms (the most important tool) showing spectral data over time with amplitude represented by darkness or color. Spectrograms reveal vowel formant structures (F1 and F2 frequencies associated with vocal tract cavity sizes), consonant characteristics including voicing, friction noise, and closures, and transitions between sounds. Vowels are classified by their first two formants, with F1 corresponding to pharyngeal cavity size and F2 to front oral cavity size, allowing construction of an acoustic vowel chart that mirrors the articulatory vowel chart but differs in methodology.