Corpus linguistics is the study of machine-readable samples of authentic spoken and written language assembled in a principled way for linguistic research, enabling researchers to explore language patterns through frequency analysis, collocation identification, and concordance examination rather than relying on intuition or prescriptive norms.
Corpus Linguistics Taster: Spot Language Patterns | UH English
Added:hi my name is Saskia Kirsten and I'm a senior lecturer and English language linguistics at the University of armature and amongst other modules I teach two that use corpus linguistics to explore language data one of these modules is called chunk language investigating for max sequences and the other one I code teach is called corpus based studies in English language and I'm going to explain what the corpus in all of those names means so in today's test a session I'm going to talk you through how we can explore language data using corpora and of course it spot the pattern because that's what we're going to do and there will be a couple of opportunities for you to have a go yourself so just follow along and then pause the video and go away do your little analysis and then come back to see what I've got to say about it okay so why patterns patterns in language occur quite a lot but we're not necessarily aware of them so for example if we look at something that lists a mock of blank the most of you will probably fill the blank with either tea or coffee there are of course other options but those are the ones that come as readily to mine so tea or coffee would be really good fits for this blank or if you're here happy blank then the most common phrase is happy birthday for example or Once Upon a blank so most of you are probably familiar with once upon a time which is a very very common phrase to introduce fairy tales for example and maybe slightly less obvious is something like blank blank ladies and what goes in here are two adjectives for example little old ladies and again it's interesting that as little old lady's not old little ladies so even the word order may play a role here I don't know whether you're familiar with Nathan pile strange planet webcomic series if you're not then go away have a look at his Instagram page so that you get an idea of what I mean by this so in these web comics there little aliens and I'm not going to explain the jokes to you about the jokes amongst other things rely on the fact that they talk similar to the way we talk but they use slightly different words not what we would you can still understand it but it doesn't sound as familiar and that's something that we would look at in one of my seminars where the humor comes from what exactly he does with language to make it still recognizable and yet slightly unusual and different from how we would say it so I've used the term corpus a couple of times obviously I need to tell you what exactly I mean by this so a corpus and the plural is corpora some one corpus two or more corpora is a collection of texts used for linguistic analyses and that is usually stored in the form of an electronic database so it's a large collection of text that is accessible by a computer and corpus texts are usually very large so we've got thousands or even millions of words and we can search them using a computer program and what's most important is that these corpus texts are not made up of the linguists or native speakers invented examples but they are authentic naturally occurring spoken and written language so a corpus therefore is a large collection of texts that is searchable and the software that we use to search a corpus is called a corpus query system or sometimes referred to as of concordance er and there quite a number of these corpus query systems around some of them you have to have to be installed on your computer some are web-based and I've listed a couple and the one that we use at the University of Hertfordshire because we have got a subscription to it is one that's called a sketch engine and I'm going to show you a couple of examples with sketch engine obviously you can't access it yourself just yet but there's also an openly available option that you can have play around with and I'm going to explain how to do that a little bit further along today's taster session so corpus linguistics again just to reiterate refers to the study of machine readable of spoken and written language samples that have been assembled in a principled way for the purpose of linguistics research so at the heart of empirically based linguistics and data-driven description of language corpus linguistics which is what we call the discipline the discipline school linguistics is concerned with language use in real context so it's language in the wild as I sometimes call it so why use a corpus it's exactly that because it's language in the wild it means that it's not what I think about language but how at language is actually used because we are not totally bad at actually reflecting on our own language use because we've got all those norms and prescriptive ideas of how language should be used they cloud our judgment of how we actually use language and we don't pay attention to what comes out of our mouth most of the time anyway at least not down to the nitty-gritty detail the corpus linguists are interested in and a corpus search can be really simple and I'm going to show you a couple of examples or it can be really complex depending or we need to do with it so it keeps pace with your questions as well which is what I like about it now here I've given you two examples of frequency lists I've just asked the computer in two different corpora so two different kinds of text collections to give me the most frequent words and rank them by it from the most frequent to the slightly less frequent and the tuned corpora I've used are the British academic spoken English corpus or Bayes and the British academic written English or Bal corpus and what I'd like you to do is to just pause the video for a second have a look at these two lists so the one that starts at the top left-hand corner and then the one at the bottom and just have a think if you can work out which one is the spoken English corpus and which one is the written English corpus so which one is which okay so hopefully you've had a workable neighbor that you have been able to work it out and you probably have quite a good idea that the top one is the British academic spoken English corpus because you may have spotted in position five the voiced pause and obviously a voiced post is something that wouldn't occur in written language normally but you may have also spotted for example in position 12 and positioned 29 but they are contraction and then in position 40 there's another one so apostrophe s as in for example he's instead of he is or in position 29 we've got the newt which is of course the knot as in don't for example and then in position 40 we've got as in I have for example so I've and not I have the bottle one is the British academic written English corpus and again what you may have spotted here is that there's a lot of punctuation that's used so full stops comma and especially opening and closing brackets which is typical for academic written English because we need a lot of brackets for referencing apart from that we also got apostrophe s but that's probably would have to look at it in a bit more detail and I strongly suspect that this is probably the genitive and not necessarily a contraction but apart from that again that is the most frequent word in both of these corpora which is not surprising at all because that is just the most common word in the English language. but that's one way in which we can explore corpora what are the most frequent words and can what can does that tell us about the underlying corpus for example what else can we do so we've already looked at the most frequent one that would be rank order we also can look at raw frequency and relative frequency so that tells us how many instances of a given word at you know given corpus and then relative frequency would be what is the percentage of the total number of words that the raw frequency represents so it would give us something like statistically speaking we would expect to encounter word X Y number of times per million wires frames up per million words we could also look at collocations so frequently co-occurring words together and we can do that using collocation candidates or words edges we can also look at the most frequent phrases of a given length that occur and that's usually referred to as an Engram sometimes also called a lexical bundle now we can also explore the frequent grammatical structures in a corpus so for example is something predominantly written in the past voice for example so this is what a concordance looks like obviously I've stripped it of all the the other surrounding bits and bobs that are there but you can't see at the moment just to give you the bare lines of concordance so what I've got down the middle in red is the word that I searched for so I conducted a query of search in a corpus in this case it was the B and C the British national corpus of the word modicum and this one down the middle is called a quick or key word in context or sometimes it's also referred to as the knowed word and then I've got context to the left and then I've got context to the left and I've got context to the right and as you may see these are not necessarily always complete sentences because that's not necessary but it gives us the words that precede the word I search for and the words that followed the word I search for and if you read the lines of concordance not left to right but top to bottom you will already be hopefully being able to spot a couple of pattern and so for example if you look to just the left of the word modicum you'll see that it's mostly oh we've got a couple of counter examples there's liver and at that in there the most instances in most instances it's a modicum and then if we look to the right we can see that in almost all instances it's a modicum of and then we can see it's almost always a molecule of something and then we could have a look in more detail at what is this something that follows for example and this already brings us to the notion of collocations and one of the very often quoted quotes here is you shall know what by the company keeps because words co-occur together more often than we would expect by chance so some words just have an affinity to each other and that's the purely frequency based definition of collocation there's also a Francia logical definition which i'm not going to go into there's something to do with them restrictions on substitutability the references are on the slides if you want to read up more on that kind of collocation and then have a look at the last slide well listed all the references that have used for this little presentation so what I've gone away and done is I've asked sketch engine to determine collocation candidates for me for a mystery word so I'm not going to reveal the mystery word I'm just going to give you the collocation candidates and their collocation candidates to one position of the right only so these are only words that immediately follow our mystery word so in the next slide you will see those words and I want you to work out what the mystery waters and remember the mystery right there word has to be to the left so it has to precede all of those collocation candidates so here you can see the collocation candidates from four different corpora and the BNC ben10 1008 which means as a web-based corpus of English from 2008 the same from 2013 and then I've got another version of this corpus from 2015 so just have a look and see if you can work out which word can precede all of these and you may have worked out that the word that we are looking for our mystery word is smart and you can also see that it develops through time so one thing that may have given it away is in the B and C you may have heard of smart aleck or smart ass and for example but then we can also see it's smart enough smart cards and then of course in 2013 and 15 we've also got the smartphone which didn't exist when the BNC was compiled incidentally and we've also got the smart TV and the Smart Grid so this these are new technologies that have come into existence and therefore also obviously have shaped the language because now all of these terms phone grid meter are very very common collocations for smart and there weren't a couple of decades ago so we can also say something about the representativeness of corpora because corpora obviously have to be compiled at one point in time which automatically means that there's a cut-off point when then there's no new data added unless it's a monitor corpus but that's something different which I won't go into so for example what I've done here is I've done searches for the work the lemma email for the dictionary form of email that includes email emails emailing emailed and all of those regardless of spelling so whether it was with a hyphen or without a hyphen and in the British national corpus as you'll see we only have 236 examples and that's not very surprising because the BNC was designed to represent a wide cross-section of British English from the later part of the 20th century and building of the corpus began in 1991 although the text date from before that as well but was completed in 1994 so that means no text new or the 1994 are in this corpus that's just logic and of course in 1994 not that many people use email on a daily basis if we compare that to the English web 2015 or en 10 10 15 that's an English web corpus downloaded by spiderling in November and December of 2015 we do it's the newest corpus that we've got on sketch engine at the moment and that gives us more than three and a half million instances and if you look at the relative frequency which is the one given in brackets you'll see that in the B and C we've got 2 per million so every million words we would expect to encounter statistically speaking email in all its variations two times yes of course it's much much more frequent nearly 200 times per million words in the English web 2015 and that just shows us that obviously the word email is much more frequently used nowadays than it was 30 odd years ago there's also scale which is a pardon version of sketch Engine 4 is mostly intended for language learners and it's freely available and that's what I'm going to use for you to have a go at for some of the activities so I've put the link to scale on the slides so just go there and have a look and I've also done another video that explains how that works so this is just a little video to show you how you can access scale and where to find it so you just go to for example all search engine of your choice then you search for scale and then you click on the link and then that brings you to this box and then in the search box you just type happy about and then you click on the magnifying glass to search and that should give you 40 examples of happy being used in context as you can see it's not ordered the way I showed you earlier but you still have the node work the word that you search for the string rather I should say because it's a phrase it's not just one word in red and then you can look at these 40 examples and on the next slide I've just copied those 40 examples and I'm going to explain in a little bit more detail what I want you to do with these examples so have a look at the examples listen to the instructions on the next slide pause the video and then come back once you've done your analysis and then I'm going to explain what we can glean from this analysis I'm here so that you can see so we've got my 40 examples and I've also made them a little bit bigger so that you can read them and I want you to go away and have a look at each of these 40 examples and determine our people here actually happy about something or are they not happy about something or can't you tell because sometimes you can't tell so for example use a smiley face when people are actually happy about something user-friendly for nice if it they're not happy about something and use nothing or an X or something for all the instances where we can't tell and once you've done that come back and I'm going to explain why I want you to do this and what kind of information we can glean from this kind of exercise okay so this is my analysis of these 40 examples and I've use a used green for all the examples where people are actually happy about something and again it slightly bigger so you can see so for example in line one he was still happy about who had been chosen so this person whoever he is is obviously happy the same as in example two I'm very happy about my results but in line three I've marked that red because not every artist is happy about the expanded role so obviously they're not happy or in line seven neither had been entirely happy about it so again here somebody's not happy in line ten I marked that in yellow because we can tell he asked at Indies whether they were happy about the upcoming operation and we don't know because we don't know the answer from this example and I've done that for the first twenty examples and also for the following 20 examples and overall you can see that there is an overall parity with a slide preference for saying that you are not happy about something I think it's something I have I should have counted it properly but I think there's like 20 examples of where somebody's not happy about something 217 examples where people are actually happy and three well we can't be quite sure so what's interesting about this is that happy about is quite often if not in the majority of cases and obviously we don't have enough data to make a definitive analysis he ought to make a definitive claim but in most examples that tent seems to be a tendency to say that you are not happy about something rather than that you are happy about something so there seems to be a slight preference for the negative with this phrase so we also have caught something in the sketch engine that's called a white sketch which is a summary of the words behavior and again yeah I've given you the definition of what a word sketch does and what it is from sketch engine itself so white sketch is a one-page summary of a words grammatical and collocation or behavior and it shows the words collocates categorized by grammatical relations such as wise that self as an object of the verb words that self is the subject of the verb words that modify the world etc and what this looks like in your version is scale so again you can go back to okay so here we've got an example of a word sketch in scale and if you have a look next to the search box we've got a couple of different options so we've got examples which is what we used before and then we've got white sketch so click on word sketch and then put donate in the search box and click on the looking-glass and little magnifying glass or whatever you want to call it and then you will see this particular screen so for donate we've got for example the subjects of tonight so for example donors donate volunteers donate businessmen donate and so and so forth the objects that aren't donated so for example proceeds Blood Land money funds acres organs you can donate kidneys more specifically dollars sums items and I'm using this the plural sometimes because it only gives me the lemma form so the dictionary form so for example organs would also be a collocation here oh we can have trophy or sperm so all of these are things that are being donated and if we have a look at this you might notice that all of these are things that have got value in some shape or form so we've obviously got money and dollars and things like that we've got what if body parts blood if you want to count as sperm and then kidneys and organs but for example you wouldn't say I'm going to donate urine for example because Uranus a bodily secretions like blood for example situation well but you know what I mean but it doesn't have any value so can't be used by someone else things that we donate are normally positive things that still have value that's something that we can glean from out from this little analysis and this has linked to something that we call semantic prosody so semantic prosody is this idea that a given word or phrase may occur frequently in the context of positive words or negative words most commonly the work on problematic prosody has focused on on negative associations and these associations are exploited by speakers to express evaluative meaning covertly so that's something that we don't do openly but by choosing certain words that normally co-occur with other negative words we're already full shadowing that will probably mean something negative for example so have a go yourself use the word sketch function on scale on scale but have a look at the verb perform and just have a look at what kind of things collocate would perform and whether you can glean anything from that ok so if you've done that then you may have come up with something like for example week you perform an operation or a surgery and this is really interesting because lots of the other examples would perform are for example you perform a role you perform in the theatre and so on and so forth and the reason why we use the word perform with surgery and operation is we also talk about an operation theater for example and I don't know whether you've ever seen that in old films for example that old operation theaters actually looked like lecture theatres we have got that sort of the seating in rows that go up higher and higher and higher so that all the students can see what happens while the operation is being performed so these are the kinds of things that tell us something about where words come from as well and how words are you and why certain verbs may go with certain nouns and so on and so forth now just to wrap up a couple of ideas what else we can do with this so there's also corpus assisted discourse studies which is really interesting and the aim of corpus assisted discourse studies is to uncover in a discourse type under study for example in new newspaper articles or in novels or in official documents and policy documents for example what we might call or rather what the authors I'm quoting here called non-obvious meaning that is meaning which is not readily available to naked-eye perusal so we can uncover hidden covert meaning again so for example how ideas about groups and people and race are constructed and discriminated ceramics Krishnam or theists and some work on that how refugees and asylum seekers are portrayed in the British press or how government officials construct their identity for example and again if you want to read up on any of those the references are all on the last slide corpora can also be used to explore literary texts in the sense of corpus stylistics you can have a look at the click Dickens project from the University of Birmingham mache in a mine back for example just some work on this and we can map the emergence of words so again Ramage krishnamurthi and Ryan 0.1 have done some work on climate change that sometimes is called global warming sometimes it's called climate change and again who uses what when and in what context is actually really interesting and now people have taken to calling it calling it a climate emergency for example so again it's about language change over time and what we associate with certain terminology as well it can obviously help us with our own writing for both native and non-native speakers I use corpora all the time if I'm not quite sure whether something sounds good or is actually used in that way it can also be used to scaffold foreign language learning which is the prime use of scale that you've already looked at so this is to provide examples for learners of English so that they get a feel for how the language is used and you can explore language in any other context imaginable so for example I've done some research on the language of twitter using corpora we've looked our policy documents with a colleague of mine from education and you can build your own corpora which is something that we've done for those two examples and it gets really interesting because you can work with language really almost at the click of a button but then of course the interpretation and the questions you ask are something that the computer cannot tell you and that's where we come in and that's what you're going to learn if you study English language and linguistics at the University of Hertfordshire so here are the promise references I hope you enjoyed this little taster session if you've got any questions don't hesitate to contact me my email address is on the very first slide and I hope to see you soon at the university far future bye
Up Next

Anke Lüdeling: Diachronic Corpora & Language Change Analysis
@idrhku
935 views•2015-02-09

American English OKAY Over Time: A Diachronic Interactional Linguistic Study
@Abralin
1.2K views•2020-07-30

Forensic Linguistics: How Language Solves Crimes | PBS
@pbsstoried
1M views•2024-01-25

Accent Expert Explains U.S. Regional Dialects | Part 1
@WIRED
9.3M views•2021-01-21
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Linguistics



















![[BK21 Eng.] Introduction to LancsBox X and Practice](https://i.ytimg.com/vi/vVIdTnDzcpU/maxresdefault.jpg)








![Analisis kolokasi 2 [LK 118]](https://i.ytimg.com/vi/HGh-1KKtb-Y/maxresdefault.jpg)









