Speech perception operates through categorical processing where listeners identify speech sounds based on specific acoustic cues rather than continuous variation; key cues include Voice Onset Time (VOT) which distinguishes voiced from voiceless consonants at approximately 30ms threshold, F2 locus transitions that reveal consonant-vowel boundaries, and frequency patterns for fricatives, with the syllable serving as the fundamental perceptual unit. Two theoretical frameworks explain this process: active theories (motor theory, analysis-by-synthesis) propose that perception involves internal speech production and articulatory knowledge, while passive theories (template matching, feature detector theory) emphasize sensory processing with stored neural patterns.
Speech Perception: Acoustic Cues and Theories of Perception
Added:Hello.
The hearing system cannot react to all features present in a soundwave. Thus, it is essential to determine what we perceive and how we perceive it. This enormously complex field is referred to as speech reception. The following questions have dominated research.
Question number one, does the speech signal contain specific perceptual or acoustic cues?
And then are there any perceptual units, any fundamental units of speech perception?
And last but not least, how can we model the process of speech perception? That is, are there any theories of speech perception?
A further important issue in speech perception which is also the province of experimental psychology is whether speech perception is a continuous or as often assumed a categorical process.
These and some minor questions will be discussed in the following. The speech signal presents us with far more information than we need in order to recognize what is being said. Still, our auditory system is able to focus our attention on just the relevant auditory features of the speech signal. features that have come to be known as perceptual or acoustic cues.
Now, the importance of these small auditory events has led to the assumption that speech perception is by and large not a continuous process but rather a phenomenon that can be described as discontinuous or categorical.
So the entire process can be characterized as categorical perception.
In other words, these cues are not perceived along a continuum but as fixed categories.
Let us exemplify such categories.
The first one is called voice onset time. Now voice onset time or in short vot is the point when vocal fold vibration starts relative to the release of closure.
It is crucial for us to discriminate between clusters such as bar and p and it is a wellestablished fact that a gradual delay of voice onset time does not lead to a differentiation between the voiceless and voiced conson consonants. Let us illustrate this. Now if voice onset time is long let's say 250 milliseconds what do we perceive bar clearly the voiced variant bar. Now if by contrast we make it extremely short let's say 10 milliseconds we perceive p again.
Okay so 10 milliseconds and the result is p. What about 50 milliseconds? Bar it is still the voiced variant and 20 milliseconds bar it's the voiceless one.
Now quite interestingly if we generate a voice onset time value of 30 milliseconds then we have trouble to identify what we hear. bar.
In other words, if vot is longer than 30 milliseconds, we hear bar. If it is shorter, the perceptual result is par.
So the voice onset time value of 13 milliseconds serves as a key factor as some sort of acoustic Q.
Here is another acoustic Q formant 2 or in short F2 transition. Now the formant pattern of vowels in isolation differs enormously from that of vowels embedded in a consonental context. If a consonant precedes a vowel then the second formant F2 seems to emerge from a particular point. The point is very high for C. It is intermediate for t and it is very low for purr.
Now this frequency region from where f_sub_2 emerges is emerges is referred to as f2 locus.
And it may be assumed that a gradual change from high to low may result in a gradual change from car to p if we generate the consonant in a vocalic context. Let's listen now. Here's a high locus car.
Clearly the result the result is car.
Now if we contrast this with a very low one car the result is pa and in the middle car.
We clearly hear ta. But what about these intermediate values?
Ta.
So if it is higher than the locus for t but lower than the locus for c we cannot identify the respective cont consonant.
Thus it seems that speech perception is sensitive to the locus of f2 and that the transition of f2 from the locus to the vowel is an important cue in the perception of speech.
Now another cue that we rely on in the perception of speech are frequency patterns. Now the frequency of certain parts of the soundwave helps to identify a large number of speech sounds.
Fricatives for example involve a partial closer which produces a turbulence in the air flow and results in a noisy a very noisy sound with spreading over a broad frequency range.
Now this friction noise and again ah this friction noise is relatively unaffected by the context in which the fricative occurs and may thus serve as a nearly invariant cue for its identification.
Having discussed the three central acoustic or perceptual cues, voice onset time, F2 transition, the locus of F2 and frequency patterns. Let us now see whether there is is a central unit on the basis of which we segment the incoming signal.
Studies into language acquisition, especially into infant speech processing, suggests that the fundamental unit of speech perception corresponds roughly to the syllable.
The central argument is the unavailability of obvious cues that facilitate the segmentation process.
Despite the absence of such cues, children are capable of acquiring their lexicon even though they have little or no information about the phological properties of the words. Hence, they must process some sort of information perhaps innate about the properties that distinguish one word from another. This information seems to be based on the syllable. Well, the syllable as you know consists of an onset, a peak and a coda.
This is the standard notation. Here is an abbreviated notation. And here are some examples. A very simple syllable man.
Man consists of the onset bilabial nasal. A vowel in the peak and the koda is another nasal consonant. If we take strength strength, we have three consonants in the onset. one vowel in the peak and two consonants in the koda. So this is the syllable.
Now how can all these findings be modeled? Well, there are several theoretical approaches. They are subsumed under the heading of speech perception theories.
All these theories have in common that first of all they agree that the ear amplifies the incoming signal and transmits it to the audiary nerve where a primary auditory analysis takes place that is filtering out of non-spech aspects and so on. Then an auditory pattern is being generated, form and patterns etc. using some sort of mental recognition device. And it is this device here, this device where the theories make different claims.
By the way, the output of all these theories is in any case, in all cases, some sort of funological form. So here, either in terms of phonemes, we might argue um or in terms of features that contains the relevant information.
We forgot an arrow here.
Let us now compare the different types of theories. There are essentially two types of theories, active and passive theories. Let's start with active theories.
Active theories assume that the process of speech perception involves some sort of internal speech production. That is the listener applies his articuly knowledge when he analyzes the incoming signal. In other words, the listener acts not only when he produces speech but also when he receives it.
Two influential active theories have emerged.
Here on the left you see the motor theory of speech perception. According to the motor theory reference to your own articuly knowledge down here is manifested via a direct comparison with articuly pattern. According to the motive theory, reference to one's own articuly knowledge is manifested via a direct comparison with one's own articuly pattern.
An alternative is the analysis by synthesis theory.
Now this theory postulates that the reference to our own articulation is established via neurally generated auditory patterns.
Now active theories can be contrasted with passive theories.
The passive group of theories of speech perception emphasize the sensory side of the perceptual process and relegate the process of speech production to a minor role. They postulate the use of stored neural patterns which may even be innate. Here are two influential passive theories.
Here to the left you find the theory of template matching.
Templates are innate recognition devices that are rudimentary at birth and are tuned as language is acquired.
Now the alternative is the feature detector theory where feature detectors are specialized neural receptors necessary for the generation of auditory patterns.
Let us summarize our observations.
There is evidence that the analysis of the spoken input signal is by and large sensory and that a system of subsegmental feature detectors is central to any theory of perception.
These feature detectors cope with specific acoustic cues in the signal.
So that's it. Thank you very much.
Thanks for your attention and see you again.
Up Next

Why Endangered Languages Matter | TEDx Talk by Linguist Mandana Seyfeddinipur
@TEDx
124.1K views•2015-11-09

Conversation Analysis: Key Concepts & Research Domains in Linguistics
@pointstoponder5186
9K views•2020-12-30

Speech Acts Overview | Pragmatics & Language Use
@oer-vlc
213.8K views•2012-09-16

Accent Expert Explains U.S. Regional Dialects | Part 1
@WIRED
9.3M views•2021-01-21
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Linguistics












![Phonology: Intro to linguistics [Video 3]](https://i.ytimg.com/vi_webp/1JYahvKUvPU/maxresdefault.webp)



![La percepción 🧠 Psicología [CICLO FREE]](https://i.ytimg.com/vi/fNUEEQqIUW4/maxresdefault.jpg)












![[44/55] Восприятие Речи | Как Мозг Понимает Слова? | Проф. Петухов (МГУ)](https://i.ytimg.com/vi/QFSKTSylbcI/hqdefault.jpg)














