Large corpora enable comprehensive study of language variation across genres, dialects, and historical periods by providing sufficient data to analyze low-frequency phenomena; for example, COCA (1 billion words) reveals genre-based patterns like the increase of 'awesome' in informal contexts, GloVe (2 billion words from 20 countries) exposes dialectal differences such as 'snuck' being more common in American English, and COHA (400 million words from 1810-2019) tracks historical changes like the semantic shift of 'gay' from meaning 'bright/happy' to 'homosexual'.
Using Large Corpora to Analyze Language Variation and Change
Added:you you so it's wonderful to be here there's a lot of things that we could talk about today in terms of corpus use corpora using corpora for language teaching and so on and I'm really interested in that topic but today I'm going to be vote focusing on just kind of this narrow issue of looking at variation I was trained as a historical linguist many many many years ago looking at historical change in Spanish and Portuguese so I've always considered myself in part a historical linguist but over time I've become increasingly interested in genre based issues and and dialectal issues as well so that's what we're going to be talking about today so these corpora are the corpora that I'll be talking about these as Tony mentioned are from English Corp English - corpora org and I'm going to be focusing today on three main corpora I'll be talking about khoka to look at genre based variation I'll start out with that and then talk about dialectal variation with globe which is about two billion words from 20 different english-speaking countries and then I'll finish up and hopefully I'll be able to spend a good half of the time talking about historical change and I'll be talking about I'll be using Khoa corpus of historical American English as well as some other corpora so size is not everything when it comes to corpora not at all not at all there's many features that make a corpus useful or not but certainly with really large corpora it allows us to look at a wide range of phenomena that we can't with corpora that are a lot smaller so for example Khoa at 400 million words is about a hundred times as big as the Brown family of corpora globe at two billion words is about a hundred times as big as the international corporate corpus of English that's not to say that Brown and ice aren't great corpora they are they're wonderful corpora I'm just saying once we have these really really large corporate now we can begin to look at some additional things and I address this issue in this article it's actually the introductory article in Doug viber and Randy repens book on Cambridge handbook that the English corpus linguistics chapter one there and it talks about the issue of size and what we can do with larger corpora just to give you a real quick example of this and then we'll jump into looking at variation there's this interesting construction in English where it's a perfect plus so we have the perfect plus the progressive plus the passive so for example he had been being watched so if you look that up for example in the British national corpus a million words it says that there are two tokens for this so two tokens not a whole lot you can do if you go to coca so this is a billion words now we move up to a few more tokens and corpora a little bit slower here than I'm expecting today we got about 41 tokens for that so we get about 20 times as much data and then if you go up to iWeb we're talking about over 600 tokens so obviously this is a low-frequency construction no question about that but the point here is that with large large corpora almost nothing at that point becomes too infrequent of a construction you know the fact that in iWeb you get 300 times as much data as in the BNC kinda shows you that size can be useful but it truly is a balancing act between corpus size and being able to look at variation web it's 14 billion words but you can't look at historical change you can't look at different countries very well you can't look at Jean recertification as well as you'd like so you know there might be this sweet spot somewhere around 1/2 billion words where it's large you you do have quite a bit of data but it's still manageable in terms of dividing up the text by genre or by dialect whatever okay so what I want to do is I want to start out talking about looking at genre based variation I'll do this fairly quickly because I think most people are pretty familiar with this type of thing then I'll move on to dialects especially historical so many of you are probably familiar with the Longman corpus the viber and others did back in the 1990's this was about 40 million words which for that time was a really good sized corpus but of course fibers 1999 book on variation in English on Rah's is you know the Bible we're looking at genre based issues in English I no one's ever going to come close to that I don't think and if you've used that book you've seen the kind of data that I'm going to be showing you here but I'm gonna go through it kind of quickly anybody so in coca at the most basic level you can go in and you can type in a word or a phrase syntactic construction then you can see its frequency across the major genres so blogs other web pages TV and movie scripts that's very very very informal spoken fiction magazine newspaper and academic okay so we're looking at the word bling here which means kind of ostentatious display of wet and shiny things jewels and stuff like that and when we click on magazine we see that this is the most common in entertainment and also men and women so this is kind of fashion as well as african-american and that's kind of what you'd expect if you're familiar with the English word bling or for example the word awesome oops doesn't help that I didn't have the full URL they're so awesome notice the frequency here in kind of these cores on rus' spoken fiction magazine newspaper academic but it's really interesting this word which is very you know has really really increased over time it's about five times as frequent as it was twenty thirty years ago awesome is the most frequent in these kind of new genres in coca blogs other webpages and TV and movie scripts the point here is that if we didn't B's John the Bichon wrists were added to coca just three or four months ago if we didn't have those we really would not be getting a full spectrum of use in English for this particular word okay so those are very simple things we're just looking at a particular word or phrase in this case so moving on to morphology here and all I'm doing the simple search says find me a word that ends in Al and is an adjective and when you take a look at this we've got national social political international medical financial and so on it's really quite striking how these al adjectives how common they are in academic compared to the other genres so any suffix prefix root that you're interested in and you want to see how it plays out across genres you can do that okay so if you're familiar with the Longman grammar and viber you're familiar with searches like this we're looking at the pass of a form of B followed by a past participle so this is the B passive and you can see that in rough terms the more formal the genre so moving towards academic the more common misconstruction I think it's really quite interesting that web Chandra's blogs and webs give academic a run for their money in terms of frequency and then kind the alter ego of the be passive is the get passive so you know got John got fired from his job as opposed to John was fired from his job and what's interesting here is how the get passive decreases in frequency over between informal and formal genres so there's kind of this seesaw plug world what's the tug of war going on between the be passive and the get pass it and get passive but bit by bit is killing off the be passive okay another example of just looking at frequency across genres is the this is the quoted of life and I mmm like I'm not gonna go out with her okay so you know you got about 3,000 tokens here from spoken drives crazy and they're like oh you can't stand me I love it's like Lord helped make the decision that's so like but and I he was like well then prove that the point here is when you take a look at the genres again TV and movies very very informal okay but actually spoken in this case for this particular construction has per million words even more tokens of this very very colloquial construction the quoted of like other things that you can do in terms of looking at genre based variation you can put in a construction and then you can say show me the frequency here by sections so here I'm looking at verb followed by possessive followed by way followed by preposition so for example made his way to the kitchen work his way through the problems and for each of the matching strings many many many many many different strings here you can see kind of the frequency of each of these in each of the eight macro genres and the first thing you should probably notice or that you'd probably notice here is just how frequent these are in fiction compared to the other genres I won - well maybe I'll do this search but this one takes a long long long long long time because we're looking at this very high frequency construction verb followed by an adverbial particle so here we're looking at things phrasal verbs okay and so it's looking at all the phrasal verbs in fiction then all the phrasal verbs in academic in the billion works of text and these are the phrasal verbs that are much more common in fiction and notice a lot of movement you know obviously a lot of informal stuff compared to the phrasal verbs in academic which are talking about argumentation and how processes are carried out and that type of thing so if you want to compare one genre to another genre what occurs here it doesn't occur here you can do that really easily so that's syntactic variation between genres but with a sufficiently large corpus we can also do interesting things in terms of Samantha's looking at the meaning of the given word in different genres so here I'm looking at kala cats noun kala cuts of chain in fiction versus academic so we got fiction on the left here and notice in fiction chain refers to a physical chain of gold chain something you're wearing around your neck for example or a chain-link fence whereas an academic chain is more metaphorical more figurative that's referring to a sequence of things okay so anyway those are all examples just very very quickly of using coca to look at genre based variation in Lexus morphology syntax semantics may come back to that a little bit later but for right now let's leave it there all right so that's coca and of course we can do the same type of thing with the B and C much smaller of course just about tenth the size of coca but the B and C genres of course extremely extremely well developed probably the best you're going to find in any largest court my corpus the people who created that 30 years ago did a great job in that respect okay let's move on to talking about globe so globe is a really useful corpus I think in terms of looking at variation between countries so globe is composed of about two billion words of data from 20 different english-speaking countries most of the we'll all of these are countries where English has at least some semi-official at least status that's why Germany and China and Japan are not in here thank so we got about two billion words and the point is we can just compare anything across those two billion words so actually let me come back to globe here so with the most basic level I mean if you're interested just in a particular word or phrase you could put in for example banjaxed which is something from Ireland which means really messed up screwed up and you can see that it's limited primarily to Ireland okay or for example eve-teasing so here what's interesting it's not one country rather it's a region so we have South Asia here India Sri Lanka Pakistan and Bangladesh and eve-teasing which means kind of sexual harassment is that term is limited primarily to that region so it might not not be just a country it might be an entire region that we're looking at this morning when a professor Rajan was speaking I noticed he talked about nook and cranny's versus nook and corners and so actually I should come back and I want to see the frequency in each of these 20 countries and he mentioned that nook and corner Nick and corners was limited primarily to India and other countries in South Asia and when you look at it yeah it really is other countries use nook and crannies but nook and corners is what is used in South Asia and even phrases things like rather more adjectives you can see how much more frequent it is in Great Britain than the other dialects I mean to my ears the sounds like English from the 1880s certainly nothing that I would use but in break Brittany still very very common so it's not just words it's also constructions looking at morphology we can use globe to do that as well so for example English has to past forms of sneak sneak and snuck so here we have Cronin sneak he they sneak in through the back window notice there's not huge differences between what might be referred to as the inner circle dialects they're those six but then when you take a look at pronoun snuck much much more common in the US and Canada than in Great Britain or Ireland that's something that kind of sounds very very American to folks who are not from the US looking at syntactic variation between dialects Khoa excuse me globe can be very useful for that as well so here I'm looking at the construction try and verb so in American English prescriptively this would be try to verb and try and verb is incorrect according to these prescriptive grammars and it's interesting that those prescriptive rules even though they were introduced almost a hundred years ago are still quite important in American and Canadian English so that try and the verb is much much less common than in for example Great Britain Australia where that prescriptive rule never really caught on the like construction again this is the quoted of like and I'm like blah blah blah so it takes a little bit of time to run on the two billion words you can see that yeah it is the most common in the US and people think of this is just being an American construction but you see that this used a little bit less than Canada a little bit less in Great Britain a little bit less in Australia in New Zealand and then you have Singapore here I mean go figure Singapore is always just it's just the most fascinating dialect where sometimes it out Americans and out brittish's even those dialects in terms of really colloquial constructions alright we talked before about semantic variation between genres in the case of chain but with the sufficiently large corpus we can look at semantic variation between dialects so in this case I'm looking at scheme and then I'm looking for adjectives that occur near scheme I could use the more recent syntax if I wanted it gives the same results notice that in Great Britain scheme just means plan program that's fairly neutral it's not positive it's not negative it's just there whereas an American English nefarious fraudulent evil socialist get-rich-quick and so on very negative Kollek it's in American English the point here is imagine you had a corpus that was one 100 the size of globe so instead of 2 billion words it was maybe just 20 million words from different dialects now all of a sudden instead of having 80 tokens or 60 tokens you might be lucky to have one token okay so when you start dealing with colic it's all you want to have a corpus that's that's big and robust in that sense if it's 1/100 this size it's just really not going to be possible to be using colic it's to look at differences between dialects the last thing I want to do really quickly in terms of variation between dialects the to look at how we can just do a very very simple search and have it provide really interesting insight into cultural issues so in this case I said look for adjectives wife and show me and compare India and actually India other parts of Asia and Africa compared to what kadru and others might call the inner circle dialects US Canada Great Britain Ireland Australia New Zealand so in kind of the developing world here we get things like chaste wife or obedient life boy we never want to try that in the United States or Britain I'm sure using those terms and also things like temporary wife or permanent wife I won't go into exactly what's going on there but really really important differences in terms of the culture of these countries compared to kind of the inner circle countries and just a simple simple search like this can help to point out some of those cultural differences okay so a little bit there on genre it's a little bit there on dialectal variation now what I want to do is spend the next 15 20 minutes or so talking about historical variation as I mentioned I was trained as a historical linguist many years ago working with Spanish and Portuguese then I moved into English so I'm really interested in historical change most of the data that I'm going to provide here comes from koa which is four hundred million words in the last two hundred years or so this will be updated go up through 2019 in about a month coca can also be used to look at very recent change the last 30 years or four so and then the now corpus can be used to look at very very recent change it's about ten billion words now and it is added to every single day about eight to ten million words of data every single day so it goes up through yesterday through June 15 and tomorrow morning at 5:00 a.m. it will include data from June 16 okay turning first to the KOA port bus so we've started in the other cases with lexical yes you know that's kind of intuitive these to understand so here we're just looking at words that have the stove the stove the stowing bestowed and you can certainly see how that's decreasing over time or for example swell as an adjective boy Jimmy that's a swell car you can see that very very common back in the 1930s it spikes in the 1930s and then decreases quite a bit after that no one would use this nowadays so any word or phrase that you want yeah real easy to look at that with Cola looking at prefixes suffixes and so on this is just kind of a simple search where I'm just saying find me ISM words weird ending in is M and show me the frequency of each of these in each of the 20 decades in the corpus and you can see that some of these have increased over time whereas others like despotism heroism and so on patriotism definitely decrease over time as far as syntax I'm just gonna give you a couple of examples here but I've done quite a bit of research looking at syntactic change with others have as well you know at the most basic level you could just look for example deconstruction need to verb and you see that you know it doesn't seem like it's that recent of a construction but it's actually used about 40 times as frequently as it was even 4050 years ago so need to verb definite increase there or for example the action or with negation do you have pre verbal negation with have in the possessive sense or do you have post verbal negation so for example pre verbal negation would be I don't have any time whereas post verbal negation would be I haven't any time so you could come in and you know this is a really you know complicated search here in terms of what you're doing but that basically means post verbal negation you see it decreasing and then the pre verbal negation I don't have the time you can see that that's increasing quite a bit and if you take a look at the chart here the increase in pre verbal negation don't have verses haven't maybe we see this beautiful s-curve here this is the kind of thing that makes historical English arts just go pitter patter when they see this this s-curve the pre verbal negation kind of you know lower levels in the 1800's then in the mid 1900s it really increases and then it kind of flattens out because it's by far the most common now in American English so pretty much any syntactic change that you want to look at Cola can be very very useful and again for lower frequency constructions you're just not going to have enough data in four million words you just not so you're gonna need a much much larger corpus like before we were talking about semantic variation chain for example for genres or scheme for dialects well we can do the same thing here historically so in this case what I'm doing is I'm saying find me call Achatz of gay and show me the frequency of each of those polychaetes by decade so when you do that you see that back in the 1800's colic hits of gay our bright happy flowers laugh because obviously gay meant happy joyful at that time and then somewhere here in the 1950s 60s and the 70s it changes its meaning for sexual orientation so lesbian rights marriage and so on and again imagine you how to corpus 1/100 the size of koa so instead of four hundred million words four million words now your token counts for these kala cuts I mean you're going to be you just not going to be able to look at changes in colic it's very well at all if you don't have a corpus that's robust like Cola and just as we did a simple search adjective wife and found some interesting cultural issues between different areas of the english-speaking world we can do the same type of thing historically here so this simple simple search adjective women okay and we're saying show me the adjectives that were used with women back in the 1830s through the nineteen tens compared to the last forty years or so so these are the older ones on the left here strong-minded women clever woman I mean the fact that they would need to even comment that shows how sexist this is because women weren't expected be clever back then and then all these things unfortunate abandoned wretched cultivated focusing on the moral fiber of these women okay they'll sound very very sexist to us as as well they should on nowadays and so we obviously get very very different politics nowadays so this very very simple search shows us some really interesting cultural changes at least in the United States a lot of people have used Khoa to look at issues like issues in science religion environment gender studies legal issues koa and coca have been used several times in arguments before the United States Supreme Court saying okay we think we know what this word means now in 2016 or 2020 but how was that word being used back in the 1870s or the 1880s when a particular law was written and a large corpus allows you to do that well okay so that's koa and that's the last 200 years if we're interested in much more recent change then we might want to use cocoa because coca in addition to allowing us to look at different genres also allows us to look at historical change since 1990 from 1990 to 2019 and this is because coca has almost exactly the same genre balance year by year by year by year this is really the only large corpus of English that maintains the same genre balance every single year over the last you know 20 30 years and so we're comparing apples to apples when we compare one year to the next so again a very basic level something like lexical like perfect storm you can see that that came out of nowhere so perfect storm in American English means all the conditions were there came together for this to happen and you wouldn't have really expected that back in the early 1990s not used at all then it really spikes in the early 2000s but now at least to my ears it has kind of this cliche feeling to it and so it started to actually decrease now in frequency so any word phrase you won't see the frequency over the last 30 years you can do that syntactic change so for example and up verbing you'll end up paying too much money you can't really see it very well here but if you look at the frequency per million words 13 15 16 19 20 22 every five year period it just keeps increasing which raises some really interesting questions about historical change and how that happens and again if you want to use collocates to compare two different sections of the corpus so in this case we're saying fine to me any colic it's near web and look in the early 1990s before the world wide web became a big deal and for example the last ten years so this is early 1990s on the left the last ten years on the wall right notice that when back then when they were talking about a web they were talking about a a web of people or relationships or even a spider web whereas nowadays of course is talking about technology so again we're using colic it's here to compare two different sections of the corpus here for 1990s versus the last ten years okay so that's coca and coca is nice because again it allows us to look at very recent change if we want to look at even more recent change we might use the now corpus so the now corpus as I mentioned it just keeps growing so April twenty twenty about two hundred and twelve million words made twenty twenty two hundred and thirty three million words June 2020 worked almost a hundred and twenty million words and by the end of June we're halfway through now end of June they'll be about two hundred and forty million words so every night about eight to ten million words of data get added to this corpus it starts in 2010 and it goes up through yesterday okay just to give you a couple of examples of how this works let's look for the word fidget spinner so this is one of those phrases that just came out of nowhere two or three years ago they were first that little toy that you spin it and you can see that it definitely spikes in the first half of 2017 and it looks like right around the middle of May so you can even see the frequency in ten des periods since 2010 so fidget spinner peak somewhere in the middle of May 2017 so there's some examples if you look in Google Trends which measures what people are searching for not how much things are used in text but how much people are searching for it when does this spike the middle of May 2017 so the now corpus meant or mirrors perfectly perfectly what people were talking about and searching at that time even to a 10-day period okay one other example of now this is something that a lot of people have looked about looked at thought about fake news okay when did fake news when did this term start it's really not being used 2011 13 15 first half of 2016 but then in 20 the latter half of 2016 it really spikes and you can even come in and you can search by 10-day period and you notice here that November 1st through 10th 2016 hardly anyone's using it and then by November 11th through 20th 2016 it explodes and it just increases so what happened right at that time obviously the US elections on November 8th 2016 that's when fake news just comes out of nowhere okay and so again the really nice thing about that corpus is it allows us to look with very very very fine grained detail at anything we're interested in I mean here I'm looking at Lexus but even new uses fun uses of suffixes like gait meaning scandal you can map those out last example I want to give just very very quickly a couple all months ago I got the idea of well if I'm already getting all this data every night for the now pork busts can - excuse me eight to ten million we're tonight why not extract out of that every night the article is dealing with coronavirus and that's about it's over the last three months it's been about forty to fifty percent of all the articles in the now corpus have dealt with corona virus so corona virus it starts in January January first twenty twenty and it goes up just like the now corpus through yesterday so any phrase you're interested in for example flatten the curve see the huge spike here back in March and then this is just kind of the most beautiful search in the world because it's actually iconic in the sense that the curve the frequency of flattened curve actually flattens out over time as people are not talking about flattening the curve as much or for example in the United States there's been this big issue about reopening you know when we should reopen stuff and you can see that in April of this year a lot of discussion of that but now that things are beginning to open up not as many people talking about that and the interesting thing about this corpus is rather than looking in just ten day periods you can actually look day by day by day by day by day the frequency of any of these words phrases and then again I clicked on June 13th so these are all from June 13th so the point is this corpus is designed to be hopefully kind of the definitive record of what's happening with coronavirus in terms of culture and society my hope is that this corpus is used for many years into the future as people are looking at what happened here in 2020 okay so I've got about 45 minutes I think that was the time that I had so again the point here looking at genre based variation dialectal variation historical variation using large corpora lots of things you view great thank you Mark so far so we we have lots of questions in the Q&A if you would like to have a look so Tony do you want me to take those or do you want to take the ones you think are interesting okay so let me let me have a look and then I'll ask you these questions let me see so there's one here from Mohammed he's asking is the globe corpus available I think he means to download yeah so all of the corpora that I've been talking about here my screen is still shared right Kony yes so all of the corpora that I've been talking about here you can come in and you can download that data so pretty much all of the two billion words just downloadable for globe yeah okay one by Iman he's asking dude do you have data about where your users come from and whether they are more language learners than researchers or the other way around yeah I mean that data all gets logged in a database but for privacy reasons even though theoretically I could look at it I don't look at it I mean um I guess I could look at you know not like person but by region what people in China or Saudi Arabia are searching looking for I haven't done that it would be interesting data uh-huh so this one by Sadat and the question is how do you interpret shall be deemed uh well I don't know let's take a look here and [Music] sounds kind of legal to me well that is interesting so yeah it's used in academic but here in just other webpages oh yeah notice copyright gov baizen a gov so we get all these government webpages here so that's definitely a legal term and that comes out pretty nicely in the corpus so this one my cat Adina she's asking mmm so you mentioned it's possible to do a morphological search in coca is it possible to expand on mechanisms behind this morphological search I wonder what is the precision of morphological parsing and cloaca so the corpus is not morphologically parsed at all so I mean you already have to have in mind a particular prefix suffix words containing a particular root but I should mention that you can also come in and you can create customized word lists so you could come in and you know these are synonyms of beautiful and you can make your own word list so if there's 270 words that you're interested in they have a particular morpheme you you're going to need to know ahead of time sometimes what those those particular words are but then you can just use all of those words as a clump as part of your search and so yeah the ability to have custom word lists can be useful for that okay this one coming from Alex the question is could you advise would you suggest which corpora to use when you're looking at legal English yes so there's a lot of corporate that I obviously didn't talk about here for a reason of time but for that one maybe the most interesting one would be the Supreme Court corpus these are opinions from the US Supreme Court and it's a hundred and thirty million words for the last 200 years or so so that might be the most useful okay this one coming from Tristan the question is can you speak about the accuracy level of morphological or part of speech tagging in your corpora and which types of searches might be more or less sensitive to potentially miss tag data yeah I mean so that's a good question that that's you know just kind of a general issue that people in corpus linguistics and computational linguistics are interested in accuracy and so on this all of my corpora well the English ones I mean I've got ones for Spanish and Portuguese as well but the English ones these are all based on the Klaus taggers from Lancaster University and I personally find it to be one of the most useful and accurate taggers of English in my mind certainly is accurate as let's say tree tagger or other well-known taggers but yeah I used the clause tagger and it's not proprietary it this is the team same tagger that was used to tag the original British national corpus you
Up Next

Corpus Linguistics for Variation Analysis | Mark Davies
@Abralin
2.3K views•2020-05-26

American English OKAY Over Time: A Diachronic Interactional Linguistic Study
@Abralin
1.2K views•2020-07-30

Forensic Linguistics: How Language Solves Crimes | PBS
@pbsstoried
1M views•2024-01-25

Accent Expert Explains U.S. Regional Dialects | Part 1
@WIRED
9.3M views•2021-01-21
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Linguistics







































