Consonants can be acoustically classified by analyzing three key parameters in spectrograms: (1) voicing status, where voiceless consonants lack fundamental frequency while voiced consonants show clear F0; (2) presence of silence portions and friction noise, where plosives exhibit silence followed by bursts and fricatives show continuous high-frequency noise; and (3) formant transitions caused by consonant-vowel boundaries, where consonants alter the resonance chamber shape and modify vowel formant patterns. This allows identification of consonant types such as plosives, fricatives, trills, nasals, and approximants within complex speech signals.
Reading Spectrograms: Consonant Classification Explained
Added:Basic concepts of acoustic phonetics, including how frequency, amplitude, and time are mapped onto a three-dimensional spectrogram representation.

Acoustic phonetics uses three main visualization tools—waveform views showing amplitude over time, frequency spectra displaying frequency-amplitude relationships at single time points, and spectrograms (the most important tool) showing spectral data over time with amplitude represented by darkness or color. Spectrograms reveal vowel formant structures (F1 and F2 frequencies associated with vocal tract cavity sizes), consonant characteristics including voicing, friction noise, and closures, and transitions between sounds. Vowels are classified by their first two formants, with F1 corresponding to pharyngeal cavity size and F2 to front oral cavity size, allowing construction of an acoustic vowel chart that mirrors the articulatory vowel chart but differs in methodology.

Acoustic phonetics studies the physical properties of speech sounds traveling through air as sound waves, focusing on analyzing acoustic signals during speech production. It examines signals through four key parameters: frequency (measured in Hertz, determining pitch), amplitude (sound intensity/height of wave), duration (time period of sound occurrence), and spectrograms (visual representations showing frequency vertically, time horizontally, and amplitude as color intensity). This field bridges articulatory phonetics (how sounds are produced) and auditory phonetics (how sounds are perceived). The vocal tract produces time-varying acoustic signals that can be analyzed using these fundamental concepts.

This section explains how spectrograms are created and used for speech analysis. The instructor describes spectrograms as graphical displays of time versus frequency versus intensity, where time is horizontal, frequency is vertical, and intensity is represented by color. The instructor explains the difference between wideband and narrowband spectrograms: wideband shows time precisely but frequency imprecisely (good for identifying voicing onsets), while narrowband shows frequency precisely but time imprecisely (good for identifying formant frequencies). The window length parameter determines this trade-off, as frequency is the inverse of period (f = 1/T), so improving precision in one domain necessarily reduces precision in the other.

A spectrogram is a 3D visual representation of sound with time on the x-axis, frequency on the y-axis, and amplitude (loudness) represented by color on the z-axis. This allows computers to recognize and store sound data, though spectrograms contain large amounts of data requiring significant computation time for analysis.

Acoustic phonetics studies the physical properties of speech sounds. Software converts acoustic signals into visual spectrograms for analysis. Spectrograms have two dimensions: the vertical axis represents frequency measured in Hertz (Hz), where one cycle of vocal fold opening and closing equals 1 Hz. The horizontal axis represents time. Each cycle consists of a trough (compression of air particles) and a peak (stretching of air particles). The brain decodes these signals through auditory receptors, matching them with stored speech patterns to achieve understanding.
The fundamentals of articulatory phonetics, specifically how consonants are classified by their place of articulation, manner of articulation, and voicing.

Articulatory phonetics studies how speech sounds are produced in the vocal tract through coordinated motor routines. Consonants differ from vowels by involving airflow construction. Linguists classify consonants using three criteria: voicing (vocal fold vibration, e.g., [z] is voiced while [s] is voiceless), place of articulation (where airflow is constricted: bilabial, labiodental, interdental, alveolar, palatal, velar, glottal), and manner of articulation (how airflow is constricted). These criteria provide a systematic framework for describing all consonant sounds.

Consonants are classified by three essential features: voicing (vocal cord vibration), place of articulation (where sound is produced in the mouth), and manner of articulation (how airflow is modified). The IPA chart organizes consonants horizontally from front to back: bilabial (lips), dental (teeth), alveolar (tooth ridge), palatal (hard palate), velar (soft palate), uvular (uvula), pharyngeal (pharynx), and glottal (space between vocal folds). Two articulators are always involved: the active articulator (moving part) contacts the passive articulator (stationary part). For example, in [p], both lips move as active articulators; in [f], the lower lip (active) contacts upper teeth (passive). Most terms only name the passive articulator since the tongue's subdivisions are complex, except for bilabial and labiodental where both parts are specified.

Articulatory phonetics studies how speech sounds are produced in the vocal tract, with consonants distinguished from vowels by air flow constriction; linguists classify consonants using three criteria: voicing (whether vocal folds vibrate, e.g., voiced /z/ vs voiceless /s/), place of articulation (where constriction occurs: bilabial, labiodental, dental, alveolar, palatal, velar, glottal), and manner of articulation (how air flows: stops, fricatives, affricates, nasals, liquids, glides); these criteria are described in the order of voicing, place, and manner.

Consonants involve the tongue touching or approaching articulatory surfaces. They are classified by place of articulation (bilabial, dental, alveolar, velar, etc.) and manner of articulation (plosive, fricative, approximant, etc.). Examples include /z/ as a voiced alveolar fricative (friction over time) and /k/ as a voiceless velar plosive (air stored then released in burst). Diacritics modify consonant symbols to indicate precise articulatory positions.

Speech sounds are produced at specific locations in the vocal tract called places of articulation: bilabial (both lips), labiodental (lips and teeth), dental/interdental (tongue between teeth), alveolar (tongue against alveolar ridge), palatal (tongue against hard palate), velar (tongue against soft palate), and glottal (vocal folds). Manner of articulation describes how airflow is obstructed: stops/plosives completely block airflow then release it; fricatives create friction without complete obstruction; affricates combine stops and fricatives; nasals direct airflow through the nasal cavity; laterals and approximants involve partial obstruction with minimal friction. The IPA chart systematically organizes consonants by these two dimensions.
An understanding of vowel acoustics, particularly the concept of formants (F1, F2, F3) and how they represent vocal tract resonances.

Vowel quality is determined by formant frequencies (F1 and F2) rather than physical vocal tract dimensions. F1 relates to the lower resonance and varies from lowest in /a/ to highest in /i/, while F2 relates to the higher resonance and varies from lowest in /u/ to highest in /i/. Shorter vocal tract tubes produce higher resonances—/e/ has short oral tube and long throat tube, /i/ has short throat tube and longer mouth tube, while /o/ has both tubes longest. These formants can be measured from any recording and synthesized to recreate intelligible vowels, though early synthesis sounded artificial. The horizontal and vertical axes of vowel space correspond to these resonating cavities.

Formants are clusters of frequencies containing information about vowels being sung or spoken. There are three vocal formants from lowest to highest: F1, F2, and F3. These vary from singer to singer and vowel to vowel but all occupy the mid-frequency range. F3 (the highest pitch formant) occurs between 2 kHz and 5 kHz, which is the range our ear canals resonate with most sensitivity. This range controls vocal clarity—the ability to perceive what someone is saying or singing. Boosting this range improves clarity; attenuating it reduces harshness or upfrontness.

Formants are the centers of resonant frequencies in a vowel sound, representing the loudest frequencies that determine how we perceive that vowel. The first formant (F1) corresponds to vowel height - higher F1 values indicate more open vowels. The second formant (F2) corresponds to vowel frontness or backness - higher F2 values indicate more fronted vowels. These two formants can be visualized on a graph where F1 is plotted on the Y-axis and F2 on the X-axis, creating the familiar IPA vowel trapezoid diagram.

Formants are resonance frequencies amplified by the vocal tract cavities (pharynx and oral cavity) that identify speech sounds; each vowel is characterized by three formant frequencies (F1, F2, F3) which can be measured and visualized using spectrograms, allowing linguists to objectively analyze and distinguish speech sounds through numerical data.

Formants (F1, F2, F3) are the resonant frequencies of the vocal tract that determine vowel quality. The source-filter model, formalized by Gunnar Fant in 1960, describes speech production as a source (vocal fold buzz) passing through a filter (vocal tract resonances). As pitch increases, harmonics spread out and formant peaks may fall between harmonics with nothing to filter. The 'soprano problem' occurs when the fundamental frequency passes above the first formant for nearly every vowel, losing the main acoustic cue for vowel identity. Female voices are less affected because their shorter vocal tracts naturally have higher formant frequencies.
The Source-Filter Theory of speech production, which distinguishes between the sound source (vocal fold vibration or turbulence) and the vocal tract filter.

The human voice is produced through the source-filter model, where the vibrating vocal folds (source) determine pitch, while the vocal tract cavities (filter) shape the sound into recognizable speech; male voices typically deepen during puberty when testosterone causes vocal fold lengthening, and voice characteristics can be artificially modified by independently adjusting pitch and formants using software like Praat.

Source filter theory explains how vocal sounds are produced. The source is the true vocal folds creating a pitch through oscillation. The filter is everything above the source—the vocal tract including the mouth, nose, and throat—that shapes the sound. In extreme vocals, there are multiple sources creating phonation simultaneously, making it difficult to study. The depth of sound can change either through filtering (shaping the vocal tract) or through the source itself (changing vocal fold behavior).

According to the source-filter theory, sound is generated at the glottis (source) and passed through the vocal tract (filter). Vocal folds create periodic waves with energy at multiples of the fundamental frequency. The vocal tract acts as a filter, transmitting more energy at resonant frequencies. Sound waves are described by fundamental frequency, amplitude, and spectrum. This theory explains how vocal tract shape shapes speech sounds.

Speech production involves two components: the source (vibrations generated by the voice box/larynx) and the filter (the chambers of the head and neck that vibrate and filter the source sound). The combination of source and filter produces consonants and vowels. This theory explains how different voices are created through variations in either component.

The source-filter model describes speech production as consisting of two components: a sound source and a resonating filter. The source generates raw sound (such as vocal cord vibration or frication), while the filter (the vocal tract) shapes this sound by amplifying or attenuating specific frequencies. This model is fundamental to understanding how speech sounds are produced and can be synthesized. Linear Predictive Coding (LPC) is based on this principle, assuming speech signals are produced by a buzzer at the end of a tube with occasional hissing and popping sounds.
Prerequisite Knowledge
- Concept 01Basic concepts of acoustic phonetics, including how frequency, amplitude, and time are mapped onto a three-dimensional spectrogram representation.
- Concept 02The fundamentals of articulatory phonetics, specifically how consonants are classified by their place of articulation, manner of articulation, and voicing.
- Concept 03An understanding of vowel acoustics, particularly the concept of formants (F1, F2, F3) and how they represent vocal tract resonances.
- Concept 04The Source-Filter Theory of speech production, which distinguishes between the sound source (vocal fold vibration or turbulence) and the vocal tract filter.
Subsequent Learning
- Step 01Analyzing acoustic coarticulation, exploring how adjacent speech sounds influence and overlap with one another in continuous speech.
- Step 02The study of sociophonetics and dialectology, using acoustic software like Praat to measure and compare phonetic variation across different populations.
- Step 03Applications in Forensic Phonetics, such as voice biometrics, speaker profiling, and analyzing low-quality audio recordings for legal evidence.
- Step 04Integrating acoustic analysis into clinical speech-language pathology to diagnose, assess, and treat speech sound disorders.
- Step 05Implementing acoustic-phonetic features in speech technology, such as acoustic modeling for Automatic Speech Recognition (ASR) systems.
Consonant Acoustics
0:03- 1
Classifies consonants by voicing, silence, and noise patterns.
- 2
Differentiates plosives, fricatives, and trills via acoustic cues.
- 3
Notes nasals and approximants show weaker formant structures.
The Lack of Invariance Problem and Articulatory Phonology
While learning to read spectrograms emphasizes identifying distinct, static acoustic cues (like formant transitions and friction noise) for consonant classification, this approach is challenged by the 'lack of invariance' problem in acoustic phonetics. In natural speech, there is no one-to-one mapping between acoustic patterns and specific phonemes due to coarticulation—the way adjacent speech sounds overlap and influence one another. Consequently, alternative theories like Articulatory Phonology and the Motor Theory of Speech Perception argue that speech is better understood and decoded through the underlying physical gestures of the vocal tract (such as lip closure or tongue position) rather than visual acoustic segments. From this perspective, relying solely on spectrogram-based classification oversimplifies the highly dynamic and context-dependent nature of human speech perception.
Analyzing acoustic coarticulation, exploring how adjacent speech sounds influence and overlap with one another in continuous speech.

Coarticulation refers to the phenomenon where speech sounds influence each other's articulatory and acoustic properties during production, manifesting through three primary mechanisms: assimilation (adjacent sounds share articulatory gestures, such as the /s/ sound becoming palatal before /j/), co-production (overlapping articulatory movements between sounds using different articulators, like lip rounding beginning during fricatives), and hypo-articulation (reduced articulatory precision in casual speech causing undershoot of ideal targets); these effects create formant transitions that provide listeners with perceptual cues about upcoming sounds, enabling the brain to interpret continuous speech despite the overlapping nature of articulatory gestures.

This section explains coarticulation and its effects on syllable boundaries. Coarticulation is the phenomenon where one segment influences adjacent segments, ensuring rapid speech transmission without information loss. It can be progressive (forward influence) or regressive (backward influence). In Spanish, coarticulation affects both consonants and vowels—for example, velar nasals may change depending on surrounding sounds. In rapid spontaneous speech, syllable boundaries become less distinct. The final consonant of one word may fuse with the initial vowel of the next word (e.g., 'mar azul' becomes 'marazul'). This demonstrates that word boundaries disappear in continuous speech, and syllable boundaries become more fluid, revealing the dynamic nature of speech production.

Co-articulation is the phenomenon where speech production involves overlapping phonemes rather than producing them sequentially. When producing a syllable like 'do,' the mouth begins forming the vowel shape before the vowel is actually produced. This means that the acoustic features of one phoneme are influenced by the phonemes that follow it. This overlapping production is much faster than spelling, where phonemes are produced one at a time. Co-articulation causes the acoustic features of phonemes to change with context, making it difficult for speech recognition systems to identify phonemes based on their acoustic features alone.

Co-articulation occurs when the final sound (often a consonant) and the initial sound (often a consonant) of adjacent words take place in the same place in the mouth or overlap. The term means the action occurs at the same time. This commonly happens with D, N, and L sounds which are stopped at the end of words and difficult to hear. These sounds all take place in the same position in the mouth: the alveolar ridge (behind the upper teeth). Examples include: good night → goodnight, bad luck → badluck.

Co-articulation is the overlapping or simultaneous production of sounds in speech sequences. It occurs everywhere in all languages—every time we speak, sounds affect each other. Different languages have different patterns of co-articulation. When producing successive sounds, if they involve separate articulators, they can work independently. However, when sounds share the same articulator, a compromise position is created. For example, in 'blood', the lips close for /b/ while the tongue prepares for /l/ simultaneously, with no gap between sounds. In 'place', the /p/ aspiration affects the /l/, making it voiceless. This simultaneous production is essential for correct pronunciation without inserting extra vowels.
The study of sociophonetics and dialectology, using acoustic software like Praat to measure and compare phonetic variation across different populations.

Praat is free, open-source software for phonetic analysis that enables spectrogram creation through audio file loading or direct microphone recording. The interface features a control panel listing recordings and a viewing/printing window. Recording requires selecting 'New' > 'Record mono sound' with a default sampling rate of 44,100 Hz. After saving recordings, users access waveforms and spectrograms via 'View and edit'. Formant measurement employs two methods: manual clicking to draw horizontal reference lines, or automated tracking via Format > Show Formants followed by Formant > Get First Formant (press F1 key). The automated method provides greater accuracy. Red dots indicating formant locations can be toggled off via Format > Show Formats if desired.

Sociophonetics is the interdisciplinary study of how speech production and perception reveal and construct personal speaker-related identities, including emotional states, social affiliations, and group memberships; this field examines how linguistic variation across dialects, gendered varieties, and ethnic varieties provides insights into language structure and social meaning, moving beyond traditional focus on standard languages to understand how speech reflects complex layers of individual and collective identity.

Praat is the most widely used software for acoustic analysis in phonetics, developed in 1991 by David van den Berg. It provides tools for recording, analyzing, and visualizing speech data. Key features include automatic speech activity detection, formant analysis, and fundamental frequency extraction. Automatic speech detection may not accurately capture word onsets for obstruents (p, b, t, d), requiring manual adjustment. For clinical populations, automatic analysis may be less reliable. The software allows extraction of duration, pitch, and intensity measurements, with Praat scripts enabling semi-automatic data extraction from annotated recordings.

Praat is a free software program that converts audio signals into visual representations (waveforms and spectrograms), enabling linguists and clinicians to measure acoustic properties like pitch, loudness, and formants to analyze speech patterns, identify voice disorders, and study phonetic variations across languages.

Acoustic phonetics examines speech transmission through sound waves and their physical properties. Quality recordings require minimal background noise, specialized external microphones, and controlled laboratory conditions. Computer-aided analysis emerged in the 1960s, enabling sophisticated research. Major software tools include Praat (spectral, formant, and pitch analysis), WaveSurfer (open-source visualization), Speech Analyzer (multi-level transcription), and SFS (University College London's research platform). Voice Onset Time (VOT) measures stop burst release timing, classifying stops into aspiration categories. Chiao and Chen's 2008 study compared Mandarin and English aspirated stops, finding Mandarin stops occupied higher and wider VOT ranges, supporting four-category classification systems over three-category models.
Applications in Forensic Phonetics, such as voice biometrics, speaker profiling, and analyzing low-quality audio recordings for legal evidence.

Forensic phonetics is the application of phonetic principles to criminal investigations, focusing on speaker identification, audio authentication, and voice analysis. Each human voice is unique due to anatomical factors (vocal tract length, oral and nasal cavity size), physiological factors (vocal fold vibration frequency - approximately 120 Hz for men, 240 Hz for women, 330 Hz for children), and environmental factors (language exposure, family background). Key applications include signal analysis (examining acoustic signatures of gunshots), speaker identification (comparing criminal speech samples with suspect samples), and speaker profiling (determining gender, age, regional background, and psychological characteristics from voice features). Challenges include voice disguise, intoxication effects, and the emergence of voice cloning technology. The field evolved from cases like the Timothy Evans case (1949-1953) and Lindbergh kidnapping (1932), which highlighted the importance of scientific voice analysis in criminal justice.
![SFU LING 100 - [10] Forensic Linguistics](https://i.ytimg.com/vi/3QHl5HfG5V4/hqdefault.jpg?sqp=-oaymwEmCOADEOgC8quKqQMa8AEB-AH-CYAC0AWKAgwIABABGHIgRSg9MA8=&rs=AOn4CLAjUAeBLnbV2sFTOFIwSaksjwFuRg)
This extensive section covers forensic phonetics, voice analysis, speaker profiling, and authorship attribution. Forensic phoneticians analyze speech and audio evidence, helping police conduct voice lineups where witnesses identify suspects based on voice recordings. Linguists analyze recordings for similar characteristics in pitch, intensity, and formant frequencies to find matching voices. The David Bain case demonstrates that prior knowledge and expectations (priming) can dramatically affect what people hear in audio recordings—when 200 participants listened to the same audio without knowing the case details, none heard the alleged confession. Pronunciation variations are explored through the 'can' vs. 'can't' case, where forensic speech scientists used spectrograph analysis to prove a doctor said 'can't' rather than 'can.' Speaker profiling analyzes voice recordings to determine sex, age, ethnicity, and geographic origin, with sex determination being most reliable (men: 60-180 Hz, women: 160-300 Hz). Authorship attribution identifies who wrote a text by analyzing linguistic features, demonstrated through the Robert Galbraith case where computational linguists revealed Galbraith was likely a pseudonym for JK Rowling. Stylistic analysis examines unique patterns in how individuals use language (idiolect), such as preferences for sentence structures with ditransitive verbs. In the JonBenet Ramsey case, forensic linguists analyzed a ransom note examining misspellings, word choices, and morphological patterns to clear the parents from suspicion. In the Jenny Nichol case, linguists compared specific linguistic patterns to prove that murderer David Hodgson had sent texts pretending to be Nichol.

Forensic Phonetics, a branch of Forensic Linguistics, applies linguistic techniques to analyze speech data in criminal investigations. It encompasses speaker profiling (identifying characteristics like sex, age, linguistic background, and social background) and speaker identification (determining if a voice belongs to a specific person). The field addresses technical challenges including quality enhancement of poor recordings, noise conversion, and authentication of recordings. Voice lineups involve comparing suspect recordings with witness descriptions. This domain has become crucial in modern criminal trials due to the increasing prevalence of audio and video evidence.

Forensic phonetics has grown significantly, using laboratory instruments to identify speakers by comparing voice samples. Speech technology works on three fronts: synthesis (computers vocalizing text), recognition (machines understanding speech), and interaction (human-machine dialogue). These systems power applications like voice-to-text in cell phones and automated telephone services. Translation professionals need knowledge of sound systems in both source and target languages to ensure accurate communication. Understanding phonetics and phonology is essential across multiple professional fields.

This comprehensive introduction covers the foundational concepts of forensic phonetics, the application of linguistic knowledge to legal contexts for voice evidence analysis. The lecture explains that voice evidence includes recorded sounds from crime scenes and suspect recordings, analyzed using sound spectrograms to create 'voice prints.' Three core applications are discussed: detecting voice manipulation, speaker profiling to identify demographic characteristics, and speaker comparison to determine if two voice samples come from the same person. The analysis reveals multiple information layers: phonemic information (speech sounds and meanings), physiological information (body size, gender, age), sociolinguistic information (social class, education, dialect), and individual-specific characteristics. The most challenging application is speaker comparison, which requires rigorous scientific methodology beyond simple listening.
Integrating acoustic analysis into clinical speech-language pathology to diagnose, assess, and treat speech sound disorders.

Alpha ratio analysis shows negative values where more negative indicates steeper spectral slope and higher dysphonia probability. Voice Report provides jitter (normal <1%) and shimmer (normal <3%) measurements. Spectrogram analysis reveals fundamental frequency (women: 190-262 Hz), intensity (60-80 dB), and formant patterns. F1 correlates with vocal opening (higher = more open), F2 with vowel anteriority (higher = more anterior). Harmonic analysis shows resonance patterns through dark/light zones. Voice 6 showed normal F0 (240 Hz) while Voice 1 exceeded normal range (396 Hz). Acoustic analysis is essential for voice disorder diagnosis and rehabilitation in speech-language pathology.

This segment covers the clinical framework for speech sound disorders: (1) Evaluation and screening using specific filters and tools to identify the true origin of phonological alterations; (2) Determining when referral or intervention on other aspects is necessary; (3) Contextual elements and specific questions for clinical assessment; (4) Clinical diagnosis using updated clinical nomenclature; (5) Qualitative and objective diagnosis documentation; (6) Clarifying the speech-language pathologist's role according to diagnosis and treatment approach; (7) Strategic treatment planning based on phonetic repertoire analysis according to patient age. The presenter also covers treatment planning: organizing an intervention route with a five-step process for adapting praxias objectively and functionally, installing correct articulatory patterns, and providing ready-to-use tools for the three stages of intervention (conscientization, production, and transfer).

A comprehensive voice assessment involves gathering patient history, conducting perceptual evaluations using tools like the GRBAS scale, performing oral motor examinations, measuring maximum phonation time and S/Z ratio, and conducting acoustic analysis using Praat software to record and analyze pitch glides, habitual pitch, intensity, and perturbation measures (jitter, shimmer, noise-to-harmonics ratio) to establish baseline objective measures for tracking voice disorder progression and treatment outcomes.

Proper diagnosis requires assessment of phonetic and phonological production, detailed history-taking, and orofacial morphological examination. The non-symptomological diagnosis (DSM-5 based) is straightforward, while symptomological diagnosis requires detailed error pattern analysis. Intervention focuses on phonetic and phonological alterations: phonetic errors use proprioceptive cues and phonetic guidance, while phonological errors require addressing cognitive-linguistic levels beyond simple discrimination tasks.

Modern voice assessment requires specific equipment: microphone with flat frequency response, external sound card, and proper recording setup. The VoxMetria 5.0 system provides user-friendly extraction of CPP, CPS, and other measures from sustained vowels and connected speech. CPP measures how well harmonic energy stands out from noise, while CPS measures the cepstral peak slope. The system allows comparison of detailed measures (CPS) versus global measures (CPP). Acoustic analysis is essential because it makes the abstract phenomenon of voice concrete and measurable, providing the best documentation under standardized conditions.
Implementing acoustic-phonetic features in speech technology, such as acoustic modeling for Automatic Speech Recognition (ASR) systems.

The acoustic model in automatic speech recognition uses Hidden Markov Models (HMM) to convert extracted speech features (such as MFCCs, delta MFCCs, and delta delta MFCCs) into sequences of hypothesized phones (the basic units of speech). The HMM consists of an emission model (probability of observing features given a phone state) and a transition model (probability of transitioning between phone states). The Viterbi algorithm is then applied to find the most likely sequence of hidden states (phones) that could have produced the observed feature vectors, enabling the system to recognize spoken language.

Acoustic models map feature vectors to phoneme sequences using probabilistic frameworks. Hidden Markov Models use states representing phonemes with transition and emission probabilities learned from training data. Deep Neural Networks provide posterior probabilities over phonemes for each frame. Pronunciation models link phoneme sequences to words using expert-derived dictionaries, though real speech exhibits extensive variation beyond dictionary pronunciations, with words averaging 4-5 pronunciations in spontaneous speech due to accents, speaking speed, and individual patterns. Articulatory phonology offers an alternative representation using multiple streams of vocal tract variables, explaining variation through asynchrony between articulatory features.

Acoustic modeling in speech recognition involves learning emission probabilities that map between hidden states (sounds) and observed acoustic features. Each sound has a characteristic distribution over frequency-domain features extracted from speech signals (typically 39-dimensional vectors). These features are extracted by taking 25ms windows of speech, applying Fourier transforms, and extracting coefficients. The model learns what acoustic patterns correspond to each phonetic unit, enabling the system to recognize which sounds are being produced from the observed acoustic signal.

Acoustic modeling is the process of converting raw speech waveforms into sequences of phonetic units (phones). The standard approach involves: (1) dividing the long speech waveform into small overlapping frames (typically 10ms), (2) extracting features from each frame, and (3) using a classifier to determine which context-dependent phone state corresponds to that frame. This transforms continuous audio signals into discrete phonetic representations that can be used for speech recognition.

Acoustic phonetics has important applications in technology and healthcare: (1) Speech recognition - acoustic features such as formants, pitch, and waveforms are critical for identifying spoken words. (2) Speaker identification - unique characteristics of a speaker's voice, including speech formants and harmonics, enable identification of individuals. (3) Text-to-speech synthesis - acoustic phonetics helps synthesize natural-sounding speech using spectrograms and formants. (4) Speech therapy - acoustic analysis is used to diagnose and treat speech disorders.
Consonant Acoustics
0:03- 1
Classifies consonants by voicing, silence, and noise patterns.
- 2
Differentiates plosives, fricatives, and trills via acoustic cues.
- 3
Notes nasals and approximants show weaker formant structures.
The Lack of Invariance Problem and Articulatory Phonology
While learning to read spectrograms emphasizes identifying distinct, static acoustic cues (like formant transitions and friction noise) for consonant classification, this approach is challenged by the 'lack of invariance' problem in acoustic phonetics. In natural speech, there is no one-to-one mapping between acoustic patterns and specific phonemes due to coarticulation—the way adjacent speech sounds overlap and influence one another. Consequently, alternative theories like Articulatory Phonology and the Motor Theory of Speech Perception argue that speech is better understood and decoded through the underlying physical gestures of the vocal tract (such as lip closure or tongue position) rather than visual acoustic segments. From this perspective, relying solely on spectrogram-based classification oversimplifies the highly dynamic and context-dependent nature of human speech perception.
in articulatory phonetics consonants are classified according to the parameters place and manner of articulation and voicing for an acoustic classification we use frequency patterns portions of silence and we know that consonants have no clear-cut forant pattern the first and most obvious parameter that allows us to identify consonants is the absence of vocal fold vibration in voiceless consonants in Assa there is no visible vocal fold vibration and thus no fundamental frequency AA by contrast involves vocal fold vibration and thus has a well defined fundamental frequency this allows us to differentiate between voiceless consonants and any other voiced sound the type of consonant whether voiceless or voiced can be classified by additional acoustic properties for example all plosives such as AA ATA and AA involve a significant portion of silence and a short portion of friction noise when they are aspirated the so-called burst the characteristic feature of fricatives such as AA Assa and Asha is the clearly marked portion of friction noise and according to the frequency range and intensity of that noise fricatives can be kept apart other consonants such as trills like AR can be identified by regular patterns of vibration and small closures in between or they even have visible forant patterns like nasals such as Ana or approximant like a yet the formance are not as clear-cut as invol s consonants influence their environment look at these two vowels A and E in isolation they have a steady formant pattern however if we put a c in between and say a key we have an additional narrowing in the vocal tract which Alters the shape of the central resonance Chambers the fings and the oral cavity so consonants considerably influence the forant patterns of the vowels with which they occur this effect is referred to as forant transition let us now identify the consonants within a more complex spectrogram clearly in this example we have at least two fricatives without a fundamental frequency they both involve highfrequency friction noise and must therefore be voiceless alviola fricatives furthermore we have three portions of Silence again without fundamental frequencies so we have three voiceless plosives the first is almost unaspirated since it involves a very short burst the second is well aspirated and number three even though it occurs finally is also aspirated which means that here we have additional friction noise in fact we have two V and one alveola plosives further we have two vowels a high vowel and a low vowel and a devoiced alveola approximate with low frequency friction noise and a nasal consonant with a clearly identifiable FN but no clear-cut form pattern the result is what I created for you a screencast
Up Next

English Sonorants: Nasals and Approximants | Phonetics Guide
@theinterestingchannel
35.4K views•2018-12-23

Conversation Analysis: Key Concepts & Research Domains in Linguistics
@pointstoponder5186
9K views•2020-12-30

Speech Acts Overview | Pragmatics & Language Use
@oer-vlc
213.8K views•2012-09-16

Accent Expert Explains U.S. Regional Dialects | Part 1
@WIRED
9.3M views•2021-01-21
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Linguistics