Large online corpora revolutionize linguistic research by providing empirical evidence for studying historical, dialectal, and genre-based variation in language. Unlike theoretical approaches that may dismiss data contradicting hypotheses, corpus linguistics values real-world language data, enabling researchers to analyze phenomena like collocations, word frequencies, syntactic patterns, and semantic shifts across different time periods, dialects, and genres. The methodology requires both sufficient corpus size (measured in billions of words) and careful attention to representativeness across different text types to draw reliable generalizations about language use and change.
Corpus Linguistics for Variation Analysis | Mark Davies
Added:oh la sesión de viento Saab rally all of you who lead missus online receive a total bang hello welcome to our early life in business online I hope you were fine I'm Maximus jean-pierre professor at federal university of region I'm glad is collaborating such a wonderful series of talks and pen sessions of this virtual event design it for the exchange of ideas within the rebirth community as you know aberlin available in viscous online is organized by the brazilian risk association incompr in cooperation with permanent International Committee of resistance latin american association of linguists in phonology Argentine society of linguistic studies Brazilian association of applied linguistics International Association of applied linguistics in whisk Society of America during whisk Society Europe in whisk Association Great Britain and Australia newest society a Berlin and these associations have a red made difference to the construction with support it has set the spark for many other wine events we thank all the Burlington for these ratings today we are going to listen to dr. Marc Davis and I will try to give you some idea of our guest speaker his expertise to create corpora and impact of his research well the best way to know his work would be to try out some of the corpora that he and his team have created dr. Marc Davis professor at linguist at the young young universe some areas of his research are corpus linguistics language challenge in genesis variation the design and optimization of the whiskey databases and threepence and bananas offer English Spanish and of stress importance to us portraits he hired this year in order to dedicate more time to research and today it was - corporate or org where it's possible to find for instance the widely used corpus of English the corpus of contemporary American English croaker which contains 1 billion words taken from 485,000 taxes related to a tremors the intelligent web-based corpse that is a large database encompassing 14 billion words from 22 meeting webpages and the new corpus coronavirus corpus which is about 304 million words in size and each continues to grow by 3 4 million words each day coronavirus corpse shows what people are actually saying you know my newspapers and magazine in 2010 english-speaking countries it is designed to be a definite record of the social cultural and economic impact of the corona violence dr. Marc Davis has written the introduction for the recent Cambridge handbook of English corpus linguistics published in 2015 and six books such as a frequent Dictionary of Spanish core vocabulary for learners offered with Kate Davis a frequent dictionary of american us works cats collocates and Connecticut lists coffered with the garden core puzzling whisk applications current sternness New Directions called edited with Stefan your eyes and Stephanie's a frequency dictionary portuguese core vocabulary for learners call for jewish anna-marie prettily to introduction improperly I wouldn't have to tell you more than what is spec at this moment and we are here to listen to professor Marc things so let's go what a mess Thank You professor Mark Davis your the title of his conference is use a large online cartridge to investigate historical dialectal and gender-based variation language sure that we are going to enjoy it I would like to invite your audience to watch the lecture and to write your questions down the chat some of the questions will be ready to our guest speaker dr. mark dates it's an honor to have your conference here thank you so much for accepting the Aberdeen's invitation so now the floor is yours thank you very very much it really is a pleasure to be here over the years I have been able to be involved with a number of scholars in Brazil and also Portugal and to work a bit with Portuguese and I have really really enjoyed those associations so when I was asked to do this I was very very pleased that really is a pleasure to be here so what I would like to do today is I realized that some people are probably quite familiar with corpus linguistics others may have a more theoretical slant when it comes to linguistics so what I want to do is spend the first maybe 20 minutes just introducing the field of corpus linguistics especially the methodology so because it's it's going to be very very different from those who then those who are coming from a very very theoretical perspective so I want spend about 20 minutes on that to start out with and then what I want to do is just to go through lots and lots of data from primarily English corpora but Portuguese corpora as well that have been created here at BYU to give some very practical and concrete examples of what I'm talking about so again first 20 so basic theory the next 30 minutes or so I will be talking about variation genre based variation dialectal variation and especially historical variation and how corpora can help us with that so to start out with in terms of the theory one thing that sets corpus linguistics apart is the fact that data really does matter it's not like we create this artificial distinction between AI language and eat language and anytime we run into data that doesn't support the theory we say oh you know what doesn't matter that's just a language nothing to see here move along no in corpus linguistics we care deeply deeply about data so it's a very very empirical field just to give one example of this about a year and a half back I had some graduate students from a very very very Theory oriented school in the United States as far as linguistics goes you might be able to guess where it is where they wanted access to lots and lots of corpus data and I was really surprised that they would contact me I thought it was just 100% theory there all the time at that school and they said well you know things are changing here recently we thought that after doing linguistics for 6070 years it might be good to actually start looking at some data and so yeah I see a bit of a change happening in linguistics where data is becoming increasingly important okay so as far as corpus linguistics goes the tools the corpora that we use the specific corpora that we use are of the utmost importance so if you're doing for example sociolinguistic variation it would be unlikely that you would spend a long time talking about the voice recorder that you used that wouldn't make any sense the actual tool the instrument is not crucial in corpus linguistics the corpora that we use are crucial so the most of the examples that I give today are going to be from English corpora org this used to be corpus BYU etu these corpora allow or provide about 100 times as much data as corpora that were available 5 10 15 years ago and you might say well so what it's that with that large amount of data it allows us to look at many many phenomena especially in terms of variation genre based dialectal variation historical variation in ways that's just were not possible 10 or 15 years ago so hopefully you can see this this is an introduction that I wrote for the for the Cambridge handbook of English corpus linguistics which came out a few years ago and in that introduction I compared corpora of different sizes and showed that with very very small corpora which is what was used 15 20 years ago we can only look at a very very narrow range of phenomena but then with much much larger court breath we can look at much more a much wider range of phenomenon just kind of following along with this especially for lower frequency constructions large corpora are very very useful so let me just give you one quick example of this Adel Goldberg who works in construction grammar her and others have looked a lot at the way construction for example verb possessive Way preposition he she wants to make her way through the crowd it's a really interesting construction in terms of the theory and what's going on and so on but the point here is that if you look in the 100 million word British national corpus there are only 16 different strings that occur at least three times or more when you go up to a billion words as in coca now you have 300 different strings so you can really begin to say something interesting about this construction and then if you were to look at the iweb corpus which has 14 billion words now you have about 17,000 unique three word strings or excuse me 17,000 unique foreword strings that occur at least three times that's very very rich data now we can begin to say something really really interesting about that so again Scott size matters one last issue and again this is something that you know if you're not in corpus linguistics you might not have thought of much about it but it's the issue of representativity or representativeness 10 15 20 years ago a lot of people who did natural language processing computational linguistics as they were trying to develop models of English or Portuguese or Chinese or whatever they would use whatever materials they had available and typically those would be newspapers newspapers are really easy to get in electronic form the problem is if all you have in your corpus is newspapers you'll end up telling us a lot about newspapers but it there's really not much insight into what's happening in spoken English or Portuguese or academic English or Portuguese if you really want to explain what's going on in a language in all its variety you need to have lots of different genres and I'll give you some examples of that in a bit so someone might say well if size is so important why not just use the web just use Google or being or your favorite search engine and just use the web which has trillions and trillions of words of data the problem with using web-based data and it can be used very very well but the problem is there's all these kinds of things that corpus linguists like to focus on colic at certain kinds of patterns syntactically oriented searches variation and so on that would be difficult and in most cases impossible with just searching by Google or Bing or whatever and I think I'll skip this one point here in the interest of time but the last thing I wanted to mention before actually get into some data is the fact that corpus linguistics is a field that it actually isn't a field it's a methodology that can provide data for people from a wide range of fields in linguistics if you're interested in syntax or morphology or semantics or dialectal variation corpus linguistics is a methodology that can provide data for you in all of those different fields corpus linguistics data is used also very very widely outside of corpus linguistics for example in or outside of linguistics for example in legal studies cultural studies and also in industry if you think about it your cell phone the data that Google or Apple uses to process your input what data is that based on that's based on corpus data so I and many others have sold lots of lots of lots and lots of data to large tech companies and that underlies a lot of what they do okay so just a little bit more on just now an introduction to some of the kinds of things that people in corpus linguistics like to do kind of the bread-and-butter methodologies that people use one spend just a bit of time on that and then we'll spend the rest of the time looking at the genre based variation dialectal variation and historical variation okay so to start out I want to go over to coca this is the 1 billion word corpus and in coca I'm going to do what's called a colic it's search or a concordance line search so in this case I just looked for the word fathom in other words to understand very well an issue and these are just 200 sample the tokens of the word fathom and then in this case I've sorted by words to the left of fathom so when you scroll through these you see all of these negative negative negative negative negative words in fact about 90% of all of the words to the left of fathom are the word not or the contraction mmm now the reason I bring this up is that corpus linguists especially people like John Sinclair have pointed out that when we're talking about words for example if you're a language teacher you might teach from a vocabulary list but words do not occur in isolation in a dictionary they might but in the real word world words are a function of the patterns in which they occur if you use fathom you're going to be using a negative word before that or it would be really strange if you found in what I'm saying dude I totally fathom what you're saying that just sounds crazy inning okay so words and constructions go together like this and we probably don't want to separate those I'm going to talk about colic it's for a little bit so in coca for example I can do a search for adjective eyes and the fact that I can search for adjectives obviously that means that the corpus is tagged for part of speech so any good corpus I shouldn't say that but most corpora are going to be tagged for part of speech and we can use that as an integral part of the search okay so we we have adjective eyes here okay that's just a very very simple example of colic it's nearby words we're looking at one word to the left of eyes but you can get a lot more complicated than this so for example let's take the word sprawl which means just spread out so you could say John was sprawled out on the floor okay if you look in dictionary comm and you look at sprawl as a noun it says the active after instance of sprawling a sprawling posture a straggling array of something that's not very useful okay if you look at Google Images and you say ok Google what do you think sprawl means this is what sprawl means visually it means a city that just keeps going on and on and on so let's look in the corpus for what kinds of words we get around sprawl so I'll look for sprawl is a noun and now I'm going to look at the colic it's so these are nouns that occur near sprawl and so we have things like you know land it keeps on going the city is growing so City and growth and development but notice that we also pick up words like pollution congestion Atlanta Los Angeles LA an abbreviation for Los Angeles those are cities that have really really uncontrolled sprawl so here's the point the fact that you get words like pollution and congestion when you look at words near sprawl see we've moved beyond just a dictionary definition to include what is said about this word in the real world when people think sprawl what do they think and so that starts getting into psycho linguistic territory that's the kind of thing that corpus linguistics ample of this I'm going to look for the word cause as a verb and these are the call Achatz that occur near cause as a verb problem damaged death pain disease harm trouble cancer injury notice how negative these are and so corpus linguists talk a lot about semantics prosity that you have this word that looks very very innocent it doesn't look like it's gonna be negative or positive but then you start looking at how it behaves in the language and you say wow that's a really negative word or sometimes a very positive work I wouldn't have guessed that simply from the dictionary definition that's why if you say in English for example that causes incredible happiness it just sounds strange because we're not used to having a positive word after it so in this case I'm looking at a 1 billion word corpus of Portuguese Court who support the gays and it's got data from Brazil 600 million words as well as Portugal Mozambique and Angola and the same thing if you look for nouns that occur after Quezada in Portuguese problema Dunn's dun mbaku more cheap disconfort okay these are really really negative words so it's kind of interesting that you get that in both languages let me give you one last example of colic it's something that's kind of interesting so in this case I'm looking for synonyms of potent that occur before a form of argument argument arguments so that's what's going on here but we're looking for synonyms so suppose that you're a learner of English and you have a thesaurus and you've learned that the word potent there's a really good word it means powerful and so on so you go around using the word potent all the time okay just because it's a synonym and you don't know what the frequency of the different synonyms are in this particular context so you can do this corpus based search synonyms of potent followed by argument arguments and you see that sure enough there is a case of potent arguments okay it occurs three times in one billion words but there are much much much better ways of saying this strong argument convincing argument powerful art of persuasion especially if you look in academic English those are going to be much much much more common now the reason I bring that up is that corpus linguistics is a crucial tool that people use for language acquisition it provides very very rich data that learners can use to acquire a very native-like understanding and use of a second language there are almost no good materials for second language learners of English nowadays that are not corpus based for Portuguese yeah still find something that are not corpus based but hopefully over time even more and more of those will become corpse basement okay one last thing very quickly and then we'll talk about variation um if we have a corpus that is composed of a billion words ten billion words or whatever this should give us some indication of the frequency of words in that particular language and so for example there is a lot of frequency data for English from word frequency dot info that is based off of coca that shows you the top 60,000 words in the language so it when you're looking at the most frequent words here these are words that you'd all know even if you're a if you don't speak English natively get down around 15,000 maybe some of those you wouldn't know down around 30,000 yeah there should be quite a few of these that you wouldn't know down around 45,000 there's gonna be lots and lots and lots of words that you don't know but the point is here that this word frequency data can again be used for materials for language learners and also for companies who are developing natural language processing tools so for example this week I'm working with I won't say the name of the company but a very very very large tech company in the US who wants to get word frequency data to use for their technology products and for Portuguese there's also uh from the corpus delicti case there is a site that you can go to where you can browse through the top forty thousand words of Portuguese okay and for each one of these words you can get lots and lots and lots of data the colic its topics that are co-occur on webpages with that word the word in context frequency information and so on so both for English and Portuguese there's a lot of that that's available there's also a book as Liz mentioned the top fifty or to me five thousand words of Portuguese based on corpora okay so that's an introduction to corpora what I want to do now is you know we could go a number of different directions from here we could talk about corpus linguistics data for language learning I've done just a teeny bit of that but that's a huge field I can talk about that for the next two hours I could talk about using corpus data to test syntactic theories I could talk about using corpus data for cultural studies but I'm gonna focus for the next twenty five thirty minutes on corpus data in terms of looking at variation and we're going to look at genre based variation dialectal variation and historical variation and there'll be a link that will be posted to this page that I'm working from and you're welcome to come in afterwards take a look at some of these searches they're searches here both for English as well as for Portuguese I probably won't have time to get through all of these but enough to give you some sense of how corpora can be used to look at variation okay so just at the most basic level in this case I'm looking I want to talk about using corporate to look at Jean reveille station so at the most basic level you can come in type in any word phrase that you want choose chart and then it will show you in the eight main genres of coca going from very informal TV and movie language from sitcoms and so on spoken language to very formal language for example academic research papers how frequent that word or phrases so this word Ling we see that it's used quite a bit in TV and movies and for those of you that don't know what bullying means it means shiny stuff that you buy that shows how rich you are okay kind of conspicuous consumption and so it says yeah TV and movies people are showing their bling a lot okay but it's also common in magazines and you can even say well what kinds of magazines does this use the most in for example entertainment famous actors and singers African American it's very common in African American culture talk about bling as well as magazines dealing with fashion for men and women but it's not really used very much in financial magazines and financial sections of newspapers okay so that's just a very very simple example of looking up a word and then seeing what types of genres and subgenres and domains a word occurs in let's take another example here I'm looking for words to end in Al that are adjectives so this short little search will bring up al adjectives and when you take a look at this the first thing that should strike you is just how common these are in academic compared to for example TV and movies okay so a lot of times people say well is is phenomenon common in language why and for me that question really doesn't even make sense okay because when we're talking about language why we're talking about all of these different genres and so it makes sense to say that for example Al adjectives yes they're very common in academic but they're really not very common at all in very informal language in TV and movies for example and that's something that a lot of linguists just don't pay attention to they'll look at syntactic constructions for example in this area well this is very common or this isn't very common that doesn't make any sense you need to say this is common or not very common in this genre compared to another genre okay let's take another example of that let's look at just the passive the be passive so here I'm looking at a form of beef followed by a past participle and this shows me that in academic English this is quite common but in TV and movies it's not very common it's about 1/4 this common as it is in academic day or for example the gift passive this is just the opposite the get passive so for example John got fired from his job ok that is much much more common in informal language like TV and movies than it is an epidemic and actually you should be able to see here that the get Casa is actually eating away at the be passive we're not to historical change yet but if you look at the frequency of the get passive it's been increasing every 5-year period since the early 1990s and so bit by bit by bit beep asses are now being expressed as get passives over time but huge huge difference between the dialects there so we've looked at Lexus words we've looked at morphology we've looked very quickly a syntax one other example of syntactic variation by genres here I'm looking at the like construction and let me give you an example of this you can see these in key words in context and it's like but it's not what I need none like well I don't want it there and I'm like yeah and I was like oh my god and I'm like wait when did this happen and so on so if you take a look at that these informal genres use the like construction much much much more than for example academic English you know a hundred two hundred times as much so again to ask is the like construction common in English that just doesn't even make any sense is it common in informal English you bet is it common in formal English no you're hardly occurs at all okay and we could do the same thing for you know Portuguese so for example the one with morphology where we are looking at the al adjectives I can do that same thing in a corpus of Portuguese that is both historical as well as forming genres and what you'll see here is that the al adjectives feliz y all event la salida y'all and so on these are much more common in academic and news than they are for example in fiction so both for lexus morphology syntax semantics I won't spend time going through all of those but same thing you know genre really really matters so I've talked a little bit about lexus morphology syntax we just give a quick example of using corpora to look at semantic variation so here what I'm doing is I'm saying look for chain as a noun look for nouns that occur within a cloud for words to the left for words to the right of chain and but here's the main point I say show me the results in fiction compared to academic so one nice thing about the corpora from English - corpora org is that the whole the corpora really really really based around being able to look at variation and being able to say what occurs in this genre or this historical comparative period compared to the other genre the other historical period it's very very easy to do so nouns near chained when you look at the nouns near chained in fiction door neck Leatherhead gold fans fingers this is a literal chain that maybe you're wearing like a necklace okay where is an academic the words that occur near chain commodity management value analysis this is a figurative chain it's not a chain you can literally touch okay so the meaning of chain is very very different depending on the genre okay so that's an example of semantic differences between genres okay so I know we're going through this quickly but that's a little bit on genre based variation corpus linguist can a lot about genre based variation let's talk a little bit about dialects okay so for these searches I'm going to be using globe which is a 2 billion word corpus based on 20 different english-speaking countries and it's about 2 billion words overall okay and so it allows us to compare words phrases syntactic and structions whatever we want in these 20 different dialects of English okay so just to start out with do some very very simple things obviously things like words a word that probably very few of you are familiar with I wasn't before I created this is banjaxed okay and when you search for this you see that it is used almost exclusively in Ireland so banjac means screwed up messed up things aren't going well okay and so it shows us the frequency in the twenty different countries much much more common in Ireland another example I'll look for eve-teasing or she was Eve he's okay you can see here that this is found primarily in southeast excuse me South Asia and eve-teasing actually means sexual harassment okay so not a very happy concept obviously but the reason I'm showing this to you is that sometimes the variation isn't just in terms of one country it's in terms of regions okay and maybe I'll give you one last example of this we can look for phrases as well so for example rather more adjectives and you see that that is much more common in Great Britain so you get things like and showed a rather more troubled mind rather more willing to engage rather more reminiscent okay very common in Great Britain to my ears as a speaker of American English it sounds very bad I would never ever ever use this it sounds very British to me okay so someone from Great Britain yeah it's fine it's part of their language but for me it's certainly not part of them in English and the corpus shows us that data really well okay so that's just some real quick examples of lexical variation by dialect but of course we can look at morphological variation syntactic variation semantic variation and more so for example in terms of morphology a lot of English verbs have variation for the past tense so that's the case with sneek okay and we can look at for example pronoun sneaked okay so so he sneaked she sneaked it sneaked and we see that there's not a huge difference between the US and Great Britain for example but when we look at pronoun snoc we see that it's much much more common in the US and Canada than it is for example in Britain Great Britain I snuck off to my father's house we snuck up on some teams and so on so morphological variation yeah very doable how about syntactic variation so for example let's go back to the like construction quoted of like and I'm like and then she's like and then I'm like and she's like so we can look in the two billion words here and we can see that yes it is the most common in the United States but in kind of scarce that fashion it's a little bit less common but still occurs in Canada then Great Britain then Ireland Australia then New Zealand and so on and of the inner circle variety of languages this construction quoted of like spreads out really really nicely here give you one other example of syntactic variation between dialects so here I'm looking at try and verb as opposed to try to verb who try and make your cases try and get our minds focused try and salvage breakfast and so on and what you see here is that this is much more common in British English than in American or Canadian English why is that because 7080 years ago some style gate guides came out for American English where they said don't use try and verb used try to verb instead and so we see the corpus data here shows us that that prescriptive rule has ended up being really important even 70 or 80 years later when we compare the dialects let's talk just briefly about using corpus data to look at semantic variation between dialects so here I'm looking for skiing as a noun I want to find adjectives that are near scheme and I say show me the adjectives in the u.s. versus Great Britain so when we look in Great Britain we see that scheme which just means plan now this is a fairly neutral word doesn't it's not necessarily positive or negative it just means plan okay but when we look in the yes the colic 'it's the nearby words are words like alleged evil fraudulent nefarious the illegal get-rich-quick schemes so in American English scheme has acquired a really negative connotation and that comes out very very nicely in just one or two seconds via that colic it's data from the corpora one last example of dialectal variation we can look at discourse issues what is being said about a particular topic in different countries so here I'm looking for adjectives wife and so wife means essentially the same thing regardless of what country it is but I'm saying show me the colic it's the nearby words of wife in the developing world in the Philippines Jamaica Nigeria compared to the developed world or what we call the inner circle for English countries where English has you know has been there for two hundred years for example so let's just compare wife in those two countries so over in you know United States Canada Great Britain and so on nothing super super interesting there but then when we look in the developing world things like existing wife temporary wife permanent wife if I won't spend the time to do this right now but if you click on that you'll see exactly what's going on culturally it's really quite interesting and things like chaste wife or obedient wife good why that just sounds bad to my American English ears sounds very very sex okay but in other countries perhaps it's not viewed as being as sexist or certainly it doesn't bother them as much and it certainly is more common to have those adjectives near wife in these countries so again a very very simple search ends up telling us some really interesting things about social and especially cultural variation okay to finish up what I want to do is I want to talk about historical variation and we might go 12 or 13 minutes on this and then we'll wrap up I was trained as a historical linguist and so I've always been very very very interested in looking at I was trained to look at syntactic variation in Spanish and Portuguese so that's what my dissertation a million years ago dealt with the causes of construction in Spanish and Portuguese so I love historical linguistics and I like using corpus data to look at historical issues okay so obviously we've been talking about looking at words between genres and dialects you can do the same thing you can put in it in any word or phrase and you can see the frequency in the corpus of historical American English which is 400 million words in these genres over the last two hundred years okay it's about a hundred times as big as any other historical corpus of English so there's a lot of things you can do here that you can't do with other historical corpora so again in this case I just look for the word bestow which means to give that it's kind of a formal word now and you can see that it's quite common 200 years ago but almost every single decade it's gone down and down and down in frequency sometimes you get this spike you the frequency of the given word so for example swell as an adjective so an example of this would be that's a swell thing to say that swell what a swell lot of merry-go-round writing it just means really really good but the point here as you can see that it spikes in the 1930s and then goes down game was really only there for 20 or 30 years historical corpus allows us to look at lexical variation really well I'll skip a few of these here let's jump to syntactic variation so for example in English okay a hundred years ago have in the possessive sense not the auxilary but possessive I haven't any idea what's going on I haven't any money a hundred years ago negation was after half but over time over the last hundred years negation has moved in front of possessive have and of course we have to use do do support okay so we could go to this 400 million word corpus Khoa and this is the pre verbal negation don't have you can see that that's increasing the post verbal negation haven't any you can see that that's decreasing and notice we can do very very very complicated searches here involving you know certain types of noun phrases and so on and so you see that the post verbal negation is decreasing but here's what I want to show you if you were to say what percent of all of the 14,000 tokens are pre verbal don't have as opposed to haven't you see this beautiful beautiful s-shaped curve where the post excuse me the pre verbal negation increases decade by decade slowly at the beginning than a huge shift between the 1920s and the 1970 and then it's kind of leveled out because that's what occurs 95% of the time now the reason I bring up this example is that in some models of historical syntax it's supposed that syntax changes very quickly because of parametric change okay that so you're gonna get a change from A to B occurring in one generation just like that but when you look at actual corpus data syntactic change tends to occur over a hundred two hundred three hundred years which raises some very very very important questions about what syntactic models can most realistically account for actual data and a lot of people who do historical syntax are gonna say that a corpus based model that's more functional something like construction grammar that almost predicts that things would take place over a hundred two hundred three hundred years there's going to be better what about syntactic change there excuse me semantic change so I want to choose a word that has undergone a lot of semantic change over the last hundred years so gay is a really really good example of that so in this case I'm saying show me colic it's nearby words words that occur around gay and show me the frequency of those words decade by decade by decade over the last 200 years so when I do that what we see here is that back in the 1800s we get works like bright happy flowers laughs because gay just meant happy joyful okay whereas somewhere here in the 1950s 1960s 1970s it underwent semantic change so now in the 1990s 2000s we get colic it's like lesbian rights marriage and so on so obviously at this point the most common meaning of gay is sexual orientation okay now here's the point and I haven't really talked about this tons to this point this is a crucial crucial issue in this case we're looking at the frequency of colic its words near gay decade by decade and take a look at the number of tokens here 1847 can and so on this is a 400 million word corpus if we had a four million word corpus four million words that was about as big as they came ten years ago okay May 2010 if I were talking about this I'd be using a four million word corpus we would have about 1/100 the number of tokens instead of 67 we might have one instead of 18 or 16 or 13 we'd probably have a zero tokens what that tells you is that using colleges to look at semantic change this would have been completely impossible a hundred or ten years ago okay and so this comes back to the issue of corpus size corpus size really really matters because it allows us to look at a wide range of phenomena that we couldn't of Khoa has kind of been a game-changer in looking at many many types of language change okay discourse change so it's interesting probably 60 70 percent of the people who use koa they're not even linguist they are cultural historians or legal scholars or scholars looking at the environment or religion or whatever and they want to see how our view of things has changed over time so in this case I just do a very simple search I just say find me adjectives before women that are common back in the 1800s compared to the last 30 or 40 years so these are the older adjectives okay noblewoman clever woman abandoned wretched pure okay devoted um you can see that back at that time women were viewed as being weak and not very intelligent very very very sexist from our point of view but that's how women were described back then whereas nowadays no we wouldn't use those things they just sound horrible to us okay so this one simple search starts getting at some really interesting cultural issues okay so that's all from koa which is four hundred million words over the last 200 years you can actually use coca corpus of contemporary American English the billion words it goes from 1990 to 2019 so 30 years you can look at the frequency of any word morphing syntactic construction whatever you want over the last 30 years using coca so in this case we're looking at the construction end up verbing they're gonna end up paying too much money and you'll see that in each five-year period it just keeps increasing in the language which raises some really interesting questions about how it is that change happens so I'll skip some of the other ones from coca but just in one or two more minutes there's some other corpora that have been created here at BYU that can look at very very very recent change so one of these is the now corpus the now corpus is composed of about 10 billion words and the interesting thing is that every single day it is updated 8 to 9 million words every day so just in the last week from May 19th to May 25th there's about 52 million words new words of data in this corpus and the nice thing about this corpus is that you can see the frequency of words in ten day increments over the last ten years so in this case I'm looking for fake news and you can see that it's very common in late 2016 but not early 2016 and in fact I can come in and I can search here in 10-day periods November 1st through 10 2016 November 11th through 20th 2016 and you see that fake news just absolutely takes off the very day after the 2016 election where Trump was elected now I'm gonna leave it to you to figure out why people started using fake news and there's a hint it wasn't Trump who used this early early on later he co-opted it okay but go in and take a look at those examples on November 8 9 10 11 2016 and you'll see some really really interesting data and then just the very very last thing that was mentioned in the introduction I've taken a subset of the now corpus and every day this coronavirus corpus is updated with about three or four million words related to that come from articles dealing with coronavirus and so just within the last five days there's 16 million new words of data May 11th through 20th they're 20 million words of data and you can search for anything dealing with coronavirus so for example flatten the curve that phrase you can see that it was used a lot back in March yet peaked in early April and flattened the curve has become more flat which is kind of interesting over time now the reason I bring up those examples is that a lot of times linguist just look at linguistic issues and they're not looking at societal or cultural issues and for me that ruins my view of language for me language is an integral part of culture and society and being able to use corpora to get at some of those cultural and societal issues for me makes the data so so much more interesting so anyway I'm done but again corpus linguistics a little bit on the methodology and the corpus data can definitely be used to look at variation between genres between dialects historical variation and provide really really interesting insight into lexical morphological syntactic semantics and discourse related issues thank you congratulations mark babies for your work it was a lecture thank you very much for showing us the potential of this collection of corpora and to our work to our our investigations in variation and also in construction Rama we we are we on the chat we have a lot of people from different parts of Brazil and we have some questions right now how you try to to make the first one is by [Music] and she asks why doesn't where the dialects of Corpus del is for you I know for at the fictional versus non fictional selection or other possible selections sometimes numbers alone are not enough so if I understand the question first of all it's dealing with Spanish rather than Portuguese and we're looking at genre based variation so fiction versus let's say nonfiction language whether or not we're looking at Lexus or syntax or semantics is that the question of my interpreting that correctly yeah so there is I haven't I didn't talk much about the Spanish corpora but just like there's the court with super gaze which has data from Brazil and Portugal Mozambique Angola they're even larger corpora of Spanish and there's a web dialects one where you can compare 20 different countries but there's also a genre based off s it's only a hundred million words but you can definitely definitely search for you know a particular word so let's say some visa so this is like sucky soo and for cheese some visa and yeah it's much much more common in fiction and here's our fiction examples so I didn't talk about Spanish and I Allah or excuse me genres but pretty much everything we were just doing for English looking at genres you can do that with the cordless telephone yard as well hey another questions from Camilla coverage from the spirit of sunburst Federer University she asks can these comparisons comparisons be done between different social media platforms boy that's a great question so my personal view is that the issue of social media so Facebook and Twitter and so on this is going to be huge in Corpus linguistics in the future it's still relatively unused at this point and the reason is it's so hard to get data from Facebook you just they don't make their data available outside of the Facebook infrastructure Twitter they make data available messiness of dad and so on I would like to see much much much more done with social media in Corpus linguistics but surprisingly there's been very little to this point there has been some but not nearly as much as I would like to see done so that's a great question another one from on a clerical is would you say that number of the data is more important the quality of data and I think another questions related to this could you comment on the relation between evidence reliable generalization and relative control over what texts are included in that about their pragmatic and discursive designs boy there's a lot of a lot of issues in one question and I think I even forgot the the first one I repeat you just say that number of the data it is more important than quality of data no not at all not at all it's it's like a pair of scissors you need many times you need both size as well as quality so you can have for example a twenty billion word corpus of a given language these are not hard to create I mean there's tools out there to create extremely extremely large corpora from web-based data but if all you're doing is pulling in billions and billions of words from the web you have no insight into what's formal what's informal dialectal variation you don't know it's just an old-fashioned word or syntactic construction or a new one and so the best corpora I think are both large and also well designed to allow us to look at the kinds of variation that I've been talking about there and I may event answered the second half of that already but remind me on the second I reminded it could you comment on the relation between evidence reliable generalizations and relative control over what texts are included in that about their brain and discursive designs yeah so that goes back I think to the issue that I was raising at the beginning that in terms of evidence and reliability if for example we have a 1 billion word corpus that is composed of just newspaper language and we want to say okay this particular syntactic construction is really frequent in Brazilian Portuguese versus this other syntactic construction so we're comparing two sincere syntactic constructions but all we're using is newspaper data which is really easy to get as well well our conclusions may be completely correct for the newspaper genre and they may be utterly and completely incorrect for a very formal genre like academic or a very informal genre like spoken so you never want to generalize outside of the types of text that you were using if you want it to represent all kinds of Portuguese all kinds of English then you need a corpus that has all of that so yeah there's really important issues there in terms of methodology two great questions another a comment and question is from one Amara she she says I had a little bit of experience with coca and corpus but I analyzed that men is that also a valid approach large numbers generally don't give us good results I think I missed half of the sentence there but I think she's talking about the issue of versus quality again that just because something big so for example let me go here to so coca for example is 1 billion words the British national corpus is about 1/10 that size does that mean that coca is 10 times as good as the BNC no of course not the the BNC is a very very well designed corpus I tried to pattern coca after the BNC but there's some things that the BNC does better than coca there's some things I think coca does better than the BNC in terms of genres just because it's bigger certainly certainly does not mean it's better again scissors size and the quality of the techs you want both those to be working together and another a question from inaudible all large corpora are so important we'll just say something about the use of social media like Twitter and others as a resource I think we addressed that in the first question and the bottom line there is that social media are crucial they're large and they have the potential I mean Facebook Twitter tend to be very informal unless you're getting a lot of ads from you know companies and stuff but so both in terms of size as well as what they're representing in formal language they're great so are these scissors are great we got both quality and size the problem is just getting the data especially from Facebook hopefully that will change in the next five or ten years another question is from is made by Diwali she would like to know whether he whether you consider corpuscle waste could also have a political point of view oh definitely definitely so as I was talking about down here with for example the adjectives near women where we see that a hundred a hundred and fifty years ago is very sexist and it's changed over time and so I believe that that shows an improvement in the way that we view women or for example LGBT issues there certainly I believe much more sensitive now than they were thirty or forty years ago okay so those are cultural issues but there's no reason that you couldn't use this data to look at political issues as well so again the now corpus ten billion words from to 2010 up to yesterday so I mean it goes it's right up to the current time if you want to use this data to look at governmental policy on the environment or education or whatever definitely definitely and for me again that's one of the most exciting things about corpus data sure linguist can you do this I'm a linguist like using this data but for me almost the more interesting more exciting use of the data is by non linguist a look at political societal trol issues thank you another question from Christina brave honest the professor at Cairo University of the engineers she asked how is possible in large database to do with different genres and different levels of formality and informality since it is possible to have different degrees of formality in the same game as for exam speech yeah so that raises a great question in terms of the word just escaping granularity okay so for example in coca if we look at what's in coca okay we there's it's divided across these eight genres and but the question is very good within spoken I mean there's different kinds of spoken there's you know worst case scenario someone is reading something off a script or the nightly news well that's not very informal at all okay where is face-to-face conversation you know that's gonna be a lot more informal and so especially in the VNC the British national corpus as far as spoken those they do a very very good job distinguishing between these levels of formality coca doesn't do that quite as much but what I did add four or five months ago when I updated coca is I added in TV and movies the subtitles from these and you can even go in in TV and movies and you can say for example I want to just look at comedies or dramas or romance or whatever now that doesn't exactly get at the issue of formal and informal but if you think about it the language of you know different kinds of movies that James Bond versus an african-american oriented movie the language is gonna be quite different and the corpora allow you to get at some of those issues but for that the British national corpus is probably the best corpus another question is are there any theoretical or it's typically issues problems with corpora based work in terms of written spoken language in comparison so I'm sorry again I think I missed half of that but okay and repeats aren't there any theoretical or aesthetically issues problems with corpora based at work in terms of written and spoken language in comparison well certainly certainly as I've tried to you know explain here there are huge differences between the genres yeah but that is precisely what corpus linguistics tries to look at some other models of language just you know they would say well these issues of genres you know that's that's issues of language external you know what we don't we don't want to do that and in language change and you know Britain versus the u.s. that's messy that we're just gonna shove that off to e language we're only interested in I'll language I language and what corpus linguistics says is no no no no you are going to miss much of the richness of language by shoving everything away and just saying that's a language and we don't want to deal with it linguistics once both language and by language oh thank you I would ask you have retired to delay to get more time to create corpora so what are the challenges to your Warnke choose linguist corpus so what are the challenges going forward well the same for me personally the same issue that everyone has its initial time I mean the corpora from you know all of these English corpora as well as the Spanish corpora as well as the Portuguese corpora unfortunately there's no team of researchers creating these corpora it's it's just one person and that's me and it takes a lot of time and so hopefully there'll be a little bit more of that not teaching and cutting back a bit on publications and we'll see how that goes thank you very much before the end of this translation I'd like to thank the audience for being with us before participation there with comments and questions I think all the Eberly organizing team members mainly son to whom I wish happy birthday hot girl and Miguel for invite me to believe today for inspiring knowledge sharing and especially at this crucial moment for offering worldwide some of the best or busy I also thank the association's working in partnership a tizzy dirt with and for promoting the free access to estate of our discussions on the most diverse thought finally and once again I would like to thank you professor Mark Davis for your talk the fascination of your work and it was a great opportunity to learn with you about all incorporated escaped variation and Culkin's you told dr. mark davis she'll tell us something before the end of the transmission yeah I just want to say again thanks to you very much or kind of hosting this and or Abilene in general I think it is so crucial that we have more cross-cultural understanding and participation and involvement and the types of things that a Darlene are doing are wonderful in that respect so it's been a great pleasure to be here thank you
Up Next

American English OKAY Over Time: A Diachronic Interactional Linguistic Study
@Abralin
1.2K views•2020-07-30

Using Large Corpora to Analyze Language Variation and Change
@laelwebinars9025
628 views•2020-06-20

Sign Language Negation: Typology and Grammaticalization
@Abralin
1K views•2021-06-09

Accent Expert Explains U.S. Regional Dialects | Part 1
@WIRED
9.3M views•2021-01-21
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Linguistics













![[한국어 학습자 말뭉치 나눔터] 한국어 학습자 말뭉치 활용을 위한 분석 도구 활용법](https://i.ytimg.com/vi/boUGPkKaGs0/maxresdefault.jpg)
























