Consonants can be identified through specific acoustic signatures: stop consonants show silent gaps and rapid amplitude rise times, fricatives have longer, smoother amplitude rises, affricates combine rapid onset with prolonged noise, nasals exhibit reduced high-frequency energy, and liquids display a distinctive low third formant; voicing is indicated by low-frequency energy (voicing bar) and periodicity in waveforms, while place of articulation is revealed through formant transitions in spectrograms.
Speech Acoustics: Acoustic Signatures of Consonants (Phonetics)
Added:in this video we're going to talk about the acoustic properties of consonants and our goal is really to think about our linguistic organizational principles that we learn from phonetics and think about those linguistic features like manner place of articulation and voicing and for each one of those try to think about the acoustic signature where you'd look in the waveform where you'd look in the spectrogram to identify the specific place or whether something's voiced or voiceless and so on as opposed to vowels consonants don't have just a simple formant structure but they have other signs that we look for and so we're gonna go through a lot of those in this little video capsule today and the goal here isn't so that you can look at a waveform or in a spectrogram and know exactly what it is that was spoken but just to start to develop an intuition to say oh this is a stop consonant and not a fricative or this definitely looks like a vowel you know things like that so let's start with just some of the basic cues so what we have here is not speech just some you know modulated noises here and what I want to draw your attention to is what I call the rise time so how long does it take for the envelope of the sound to go from 0 up to full volume and so what this does is give us a cue for how quickly the sound bursts out to be full full intensity so this will come back as we look at different manners of articulation another thing we see here is different modulations so over on the left side here we have a slower modulation rate and then that gradually speeds up over the course of the sound another thing about modulation is that in addition to its speed it can have different depths so again on the left here we see modulations that go completely from full volume down to zero but over toward the right those modulations are pretty shallow so just to get a sense of what these sound like let's put some modulated sounds together we can listen to these here so that's a modulation rate of four per second now here's six per second a little bit faster and now twice that is 12 per second this will be pretty fast and even faster that would be 20 per second so those are different modulation rates and then we can also change modulation depth go here we can open that sound here's full depth here's where some of that depth is made shallower and as well as the sound of my phone hitting the floor there we go and here's really much more shallow depth and over on the right is depth that's so shallow you might not even hear the modulation anymore so what we want to do is start to think about how to connect those acoustic properties back to the linguistic properties so in general if you have a lot of vocal tract constriction you're going to have greater modulation depth because you're actually completely stopping the air flow and then restarting it so it should go all the way down to zero and then back up to full intensity similarly if you hold back the airflow and build up more and more pressure then once you release it it should come out with greater speed and so maybe you have a quicker rise time so let's look at how that appears in actual words so here we have the word bust and there are a few things I want to point your attention to so first we have some areas of the sound where there's really pretty much just silence and so this is obviously more visible if there's some sound before and after the silence as we can hear see here between the s and the T release before the B it's not really clear whether there's actual silence or whether the person just you know didn't start speaking yet so let's look at this as it appears in word a medial position so here we have a guy and a car okay and we can see that that that silent gap is really part of how the sound is produced and just to demonstrate that a little bit here we can look at a word like risk this here risk and then if I make a copy of that sound and actually take out that gap what we see would be this we can actually take that gap select it just as you'd select like words on a document and cut it and here's what that would sound like risk risk so you can still maybe understand the word but it doesn't sound natural as it would as the original recording with that actual gap in it risk write risk and here's the one with no gap risk so now it just sounds like a little noise add it to the end of it so those gaps are actually part of what it means to produce that sound here's another example of the word Shaq Shaq and then without the gap it sounds like this Shaq Shaq like it's been improperly spliced okay so back to our stop sounds here in addition to that gap we also have evidence of a really rapid onset which I'm highlighting here with our red arrows so that's really a signature of a stop because we're holding the airflow back and then releasing it all at once so we can look for a bunch of different step properties using praat as a tool right so we can look at these stop sounds here here's a woman producing a few stop sounds so for all of these we can see there's a gap in between the two ah sounds so in-between a baaad there's that little gap there and then in between ah and PAH there's a little gap there so that's one of them in addition if we look at the D sound here we can see that this onset is really almost instantaneous in fact we'd have to zoom way into it to see that it's it's a basically a perfectly vertical line which is one of the signatures of being a stop sound for those of you speak languages other than English particularly if you speak an Arabic language you'll notice that we have another one of those even before the first ah sound which is a signature in English that we do when we have a vowel onset word is we actually put what's called a glottal stop here which sounds like this that little a sound is how we signet signify to the listener that the word begins with a vowel sound so we have a bunch of stop sounds on the screen and let we can explore some properties such as the duration of how long we have not just after the gap but during this actual part of the sound that we can hear hear that that part is a periodic and we can count how much time has elapsed between the start of that and then the start of the vowel so that's a distinctive signature of stop sounds as well that we call voice onset time we'll have a whole video about this in the next capsule so I won't spend too much time on it except just to point your attention to the fact that we'll actually want to be keeping track of some of these landmarks like when does that burst happen and when does the valve again this voice onset time is a really well-known and well-understood property of stop sounds that if you're going to understand any consonant feature this is usually the one you start with so we've been talking about stop sounds a lot but there are lots of other consonants as well that we want to give some attention to so now we're going to move on just to think about some other properties and to do so we want to start with some of the basic properties of acoustics like the envelope the fine structure and so on so you'll remember that different manners of articulation correspond to different amounts of airflow and so because of that the different manners that we see here so we have a stop sound a fricative an affricate another stop that's voiced a lateral and then a nasal sound all of these have different ways that the airflow is shaped as its as it's coming out of the mouth and so therefore they have different envelopes that correspond to those manners so here that was what we're getting into the differences in rise time so here we have a fricative that has a longer and smoother amplitude rise time compared to for example a stop set so if we compare taw to saw what we see in the envelope of taw is it has this silent period the zero mark and then it immediately has a vertical line getting you know gaining some volume right there but for the saw it really gradually rises up to that volume so that's one of the main differences between a stop sound and a fricative sound in Africa is a bit of a combination of both where you see a rapid onset but then you see this prolonged period of noise that you can see in the spectrogram here one of the things we think about is the differences between fricatives and affricates so the main thing to think about well there are two main things one is that the affricate will be shorter in duration and this is one of the signatures of it being kind of like a stop sound stop sounds are pretty short as well as we have a more rapid onset time for Chuck than we do for sure so one of the things you can do is play with a recording of a sound such as chip okay I think we have that here well we have the word ship and so what we can do with this word ship is if we take away some of the some of the duration of the fricative sounds I'm just gonna cut this away now listen to what this sounds like chip yep now it just sounds like chip so it's really the same sound I've just blocked some of it out and so this longer duration and longer onset of the chef's sound if we just make that shorter duration with a rapid onset we can just turn it into a chess sound okay so we can move on to another category of nasal sounds and these are pretty tricky because they ask us to really think in some detail about both the waveform and the spectrogram so let's look at a word an example word here like the word name we have a nasal both of the beginning and the end of the word and what I'm pointing at are the boundaries between where the nasal in this case where the nasal ends the vowel begins and here where the vowel ends and the nasal begins so we can see these cues in the waveform that will zoom into in just a bit and what we want to think about is the lack of upper formant energy what I mean by that is something that's visible on the spectrogram only it's hard to see this on the waveform but we'll touch on it so we can see in this first area where the nasal is we really only have low frequency energy whereas for the vowel for the a in name we have a lot of this high frequency energy that comes in and as soon as that val goes away and turns into the M it kind of disappears so if we zoom into this part what we see is even on the waveform we see a lot of extra detail in this high frequency fluctuation that corresponds to that high frequency energy in the spectrogram that all goes away once it turns into the nasal and now we just have this very smooth looking periodic pattern that really just corresponds to that mmm that low frequency murmur of the nasal sound if you're wondering why that happens we can think about what happens to the airflow so for an aural sound right the air just sort of comes out of the oral cavity there's a lot of high frequencies that are there but for the nasal sound over there on the left within the oral cavity some of those high frequencies in that acoustic energy is just actually trapped in the mouth because if you produce an EM you're actually holding a closed resonating chamber in your mouth but the rest of the airflow just sort of flows out smoothly through the nose so for those of you who are interested in some of these extra details what we're talking about our knit nasal anti resonances it's a little bit beyond the scope of this class right now but this is just there for those of you who are really interested in this aspect of speech acoustics some more examples here of some nasals where we have an hazal part here at the beginning of numb then our vowel then another nasal part what we can see is that the nasal is lower energy overall because some of that energy is being stolen away again by the oral part of the vocal tract we can see the same pattern and nudge where that nasal is there but it's not quite as intense as the vowel then this part looks a little bit like a nasal but what we want to look at here is that this is so dramatically lower in intensity that it's actually just part of the gel so this part before the jaw and a nasal are very very similar and it's not really the scope of this class to always tell those things apart but it's one of the things that we can notice so the point from the slide nasals have less high-frequency energy and vowels have high frequency energy and multiple formants that we can see in the spectrogram another thing you'll notice about this is that the nasals have different formant patterns here and what this corresponds to are different shapes of the vocal check so just as we are changing the shape of say bottle as we're making resonant frequencies if we change the different shapes of the vocal tract we'll see different patterns in the frequencies so a cartoon sort of schema Thais version of this we can see here for the different places of articulation but and gah what we would see on a spectrogram are different patterns of where the highest energy frequencies would fall so what we're looking at are the transitions into those horizontal lines the horizontal lines all have different places but what we want to look at is how they actually emerge so for the B sound all of those lines begin with an upstroke okay for the D all the first ones begin with an upstroke but the second formant begins sort of at a target frequency and depending on where it goes it can either rise into it or it can fall down into it and the third third formant frequency always slides down into its position okay so for the G sound we're actually beginning with a pinched second and third formant and then after that it sort of goes into its steady state so it's a little maybe easier to look at this on a cartoon version like this but we actually do see these patterns in real spectrograms so just to give you a taste of this if we look at voiced stops like these you can look at the spectrogram directly or you can tell the computer to show you the formants and now we can see if we set the formants properly which I'll try to do here there we go then we can see these patterns where the distinct pattern of that formant tracking in red really tells us that these these sounds have different places of articulation so notice though that if I turn the spectrogram off these sounds are really hard to tell apart because on the waveform it's very hard to tell place of articulation I would say it's nearly impossible so from the time information on the waveform that gives us information about manner of articulation and voicing but to really understand place of articulation what we're understanding as changes in resonant frequencies and if it's frequency you want to look at you want to see the spectrogram stops aren't the only sounds that have place of articulation we can also see that in fricatives as well so fresh compared to X sound we can see on the spectrogram that sure has energy at a lower frequency and s has energy at a higher frequency so even if these weren't labeled if you see differences in the frequency content you can be pretty sure that where we're looking at is a change in place of articulation so let's go back to voicing just for a moment as we're looking at different fricative sounds we have a clear difference between shut and s here the F and the theta sounds are so acoustically similar that it's very difficult to tell these apart and for those of you who are interested in I can you know gladly field those questions later but we won't really focus on that so much right now what I want to draw your attention to is if we compare this voiceless series of shapes so with the corresponding voiced series of G's what we can see are a few differences in the spectrogram so I'm going to toggle back and forth just for a moment and then draw your attention to the fact that for the voiced sounds we have what's called a voicing bar down here which is just a bar of energy that corresponds to the lowest frequencies that just really tells us that the sound is voiced to compare them side-by-side we have an S on the left a Z on the right and so there are two differences I want to draw your attention to here the Z has a voicing bar down at the bottom and C also has periodicity in the wave form as opposed to the random fluctuations in s the Z has a more regular set of fluctuations that correspond to the rate of vocal fold vibration the last set of sins that we'll talk about are the liquid sounds so a full description of this is so complicated that really goes beyond the scope of an introductory class but what the one thing I want you to know about this is that there's a particular acoustic signature for our sounds which is a third formant so let's go back to this real quick here so what we have is the first formant the second formant the third formant which is the one that has this little hook in it and then the fourth format way up there so if we compare rah to LA they're very similar except for the fact that the third formant has a different shape and when that third formant sort of sinks down really low that's the signature of it being an R sound so highlighting it over on the right for lead and then on the left for red we can see how that third form it starts at a much lower frequency and that's really the signature of the R in the in the spectrogram they're going back here we can also just remind ourselves that if all we had access to were the wave forms this would not be a possible task you actually really need to see the spectrogram to appreciate that form and difference there so consonants can be are colored and vowels can be are colored as well so if we have the vowel in the name Bert this is what the spectrogram would look like and it looks like it might only have one two formant frequencies but we're actually looking at our three formant frequencies where the third one sinks so low that it gets very close to the second format so if you talk to Bert you should comment on how low his third formant is and how it makes his name our colored so just to recap here voicing and stop consonants corresponds to voice onset time which we can measure on the wave form we'll have a whole video about that in the next segment voicing for other sounds what we're looking at is low frequency energy periodicity and durational cues manner of articulation corresponds the changes in the amplitude envelope with particularly fast rise times for stop sounds place of articulation corresponds to different frequency patterns and a consonant feature of vorticity corresponds to a low F 3 there's a good website here which walks you along different waveforms and speech patterns and one of the things I like about this is that it reminds you that as you're looking for specific patterns you're limited by the exact graph that you're looking at so for example as we've pointed out a couple times in this video waveforms can tell you that you're looking at a specific kind of sound like a vowel or a stop sound but they can't reliably tell you exactly which vowel you're seeing or exactly which stop consonant you're seeing to get that exact knowledge you have to look at multiple properties all at once do some comparisons and also remember that speech is variable person-to-person and you have to be pretty experienced to get that level of detail I'm not looking for any person in the class to demonstrate that level of expertise at this time but just to demonstrate that we know that if we're looking for feature X we look for property Y on the spectrogram or on the waveform for example okay so for the next for the next set of lecture slides it'll be in the next capsule here we'll just focus specifically on voice onset time that's the end for consonant features for today
Up Next

Reading Spectrograms: Consonant Classification Explained
@oer-vlc
50.5K views•2015-07-02

Conversation Analysis: Key Concepts & Research Domains in Linguistics
@pointstoponder5186
9K views•2020-12-30

Forensic Linguistics: How Language Solves Crimes | PBS
@pbsstoried
1M views•2024-01-25

Accent Expert Explains U.S. Regional Dialects | Part 1
@WIRED
9.3M views•2021-01-21
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Linguistics












































