This video tutorial demonstrates fundamental search techniques for English corpora (COCA, COHA, NOW) on English-corpora.org, including searching for word frequencies by genre/register, using part-of-speech tags (POS) with uppercase for full words and lowercase for specific word forms, searching for lemmas (base word forms) using capital letters, finding words following or preceding other words using asterisks, and using the pipe operator for multiple conditions. The instructor shows how to switch between corpora, interpret frequency charts, and analyze linguistic patterns across different time periods and genres.
Corpus Search Tutorial: English-Corpora.org Queries
Added:well hello there students let's take a look at this activity in this activity the students will gain some familiarity with the search interface of the English corp.org G's corpora so you're going to work through alone or with the neighbor some of the um some of these these queries down here so let's take a look here uh number one frequency of the word rad in coka coka being the Corpus of Contemporary American English that c is important it distinguishes it from the Corpus of historical American English which we'll see down below so the frequency of R in Coca so you go to English corp.org right you log in I have already done so so I should be good to go and then you click on Coca right there and then you just search for rad and boom there you have your answer 1,252 that is number one number two the frequency of burst in academic spoken and magazine genres in coka so coka has eight registers so let's look at burst and you could do this one of there's several ways to do this um probably the easiest way is to click on chart that second option right there and then click go see frequency by U section so the let's jump back over the the question was academic spoken and magazines in Coca so academic is right there 88 occurrences there with a per million frequency of 7.41 and that's what that graph is representing the per frequency the per million frequency there spoken is right below it and then the magazines are much higher at 21 uh per million 2.36 per million okay so that's one way let me jump back here and show you another way down here on the bottom uh below the search form there's sections and you can highlight the exact sections you want so if we said spoken magazine I'm holding down um the command button because this is a Mac you probably hold down control and windows right and click those and then if you were to um click C frequency B section right here the search button right there it just eliminates the other ones so you can just see the ones that you specified in sections there spoken magazine and academic same numbers right these are the same numbers we saw a second ago but they're just uh they're just kind of focusing on those those right there we have zeros and everything okay so that's that one how about all adjectives in the academic and spoken registers from coka all adjectives all right let's go back to our search form click on the search up here and then you can click reset to kind of reset the search form I clicked reset right there and then um this POS right here it says POS question mark if you click on this it gives you P part of speech tags p is part of speech tags now we need to stop here for a second and discuss there are two options here right here it has capital P capital O and capital S this means that if you clicked on this and and you grabbed for example a DJ all that's adjectives it's going to put over into my search bar with a little bit of JavaScript um capital A DJ because that is the part of speech tag for adjectives okay now I just want to erase that let me just reset the whole form again with that button there that button I'm going to click on POS again now this there's a little link here it's pretty small it's not super apparent but there's undor posos right there if I click on that you'll see that the capital POS in this bar over here will change Ready Set click it now has underscore posos now what this means is that if you want to have if you want to give a word and then say I only want words that have this part of speech you would use this option so let's say just as an example real quick if I say happy and then I want to make sure that happy is being seen as an adjective um I first make sure that I have a underscore lowercase POS right there okay if it were not I'll click on this link here this switches between those two options okay again let me just review with capital P like it is right now this means if you come in here and click on this it's going to put an uppercase adj over there but if I have it this way and I click on adjective it's going to put an underscore lowercase j okay so what I'm saying is I want to find happy no spaces underscore lowercase j which means I wanted to find a happy as an adjective okay so it's a small little detail there but it's it's important to know there so let's go back to our question frequency of all adjectives in the academic and spoken registers from coka okay so we need all adjectives from those two genres or or registers so I'm going to come back here I'm going to click reset there that button says reset just to clear it out and then I'm going to select sections there I'm going to say I want academic oops I want academic I'm going to hold down controller command and click spoken let me just double check I'm going to flip back over the slide yes academic and spoken genres and Coca awesome now I'm going to use this POS tag and I can see that it's uppercase POS and so this will give me the full word so I'll click on adj all there it puts in capital adj now you could just write in capital adj and once you get familiar with these par of speech tags you can just simply manually enter those and then click on um actually I'm going to click on chart right here instead of list it's selected as list right now I'm going to click on chart so that we can see fre see frequency by section click on the button to go and what do we see here we see a lot of adjectives apparently in academic more adjectives in academic than there are in spoken hm interesting about twice as many if we look at the normalized count per million right there 96 million versus 53 million or that's that's not million 96,000 versus 53,000 there we go okay okay that's that one let's go to the next one frequency of words that end in gate in each decade from koha I I bolded the H of COA to Signal we're now going to a different Corpus within the suite of corpora under the English Dash corp.org website so the way to go from one Corpus to another you can click on this diagonal right upper right um Arrow there click and then come down here and find koha which is the Corpus of historical American English not Contemporary American English but historical now the interface is identical you can see the interface didn't change much at all except for the word contemporary to historical so just be aware of what Corpus you're in uh within this Su of of corpora okay let's go back here and look at our actual question again frequency of words that end in gate in each decade from koha so we want to find any number of word any number of characters followed by G A okay that's what we want to find there um and let me just sorry just keep jumping back and forth getting you a little seasick back and forth there sorry uh any word find the frequency of any word that end Eng gate in each decade so instead of list right there I'm going to click on chart there and then click on C frequency B section which is decade all right so here we have from the 1820s that decade all the way down to 2010s that decade and there's the frequency of gate and unsurprisingly in the 1970s in American English we saw um words that end in Gates a lot I imagine that's driven by Watergate let's just take a quick oh wait the word gate as well yeah is also a word that ends in g a t that's how you can see the frequency across the decades just so you can understand um the way to read this chart is that the freak is simply the raw frequency just how many times it occur the freaks M or excuse me the words M right there is the number of words in millions in this section so the 1830s section has 13.7 million words in it so if you take this number 1741 divided by 13.7 million you end up with the normalized frequency uh per million of 126.
N8 and these bar plots here um are graphing up the per million counts not the rock counts okay good so we got that done mhm let's look at the frequency number five the frequency of nuclear or nuclear or nuclear have you pronounce it as in as an adjective from 1940 to 1990 and koha okay so we have this word and it has to be an adjective from those um 50 years 60 years so let's go back here let's reset I'm going click our reset right there reset my search form I'm going to need uh new nuer I'm now super conscious um of that word okay I'm going to click on posos I'm going to switch it to the underscore lowercase POS by clicking this little link right there see it click now it says that I'm going to click on it and then go down to AJ and it puts an underscore lowercase j next to my word that I put in I if you can see that let me Zoom way the heck in so you can see that super clearly see that underscore lowercase j put on to the right side now again you can just simply type that in once you get familiar with these you can just simply type those in you don't have to use that drop down menu and then we need those certain decades so we need from the 1940s zo up Scroll up a little bit 1940s to 1990s I'm going to hold down shift on my keyboard shift click to select 1940s to 1990 90s those 60 years okay now I'm going to jump back here real quick this the frequency of nuclear has that from that okay we're going to click on chart so we can see a nice plot across those decades in here let's take a look from the 1940s to the 1990s we see um nuclear going up here and going down there going up in the 1960s and then a little bit down 70s back up in 80s and then down again 1990s yeah that makes sense to me just kind of thinking about the historical political events of those decades right with uh the Cold War and and nuclear arms race and stuff like that okay let's go to frequency of LMA break as a verb um I don't specify so I think we're still in the ca let's stay in the ca okay so Lemma let's yeah we're going to look for lemas and the way to search for lemas let's go back to our search form click reset there to clear out the search form the way to search for a Lemma is to use capital letters so I'm just going to type break as capital letters to find all forms of break and then I'm going to jump back to the instructions real quick we need it to be a verb and so again we're going to use this POS tag but again we need to switch it from uppercase POS toore lowercase POS by clicking this little link right there it's pretty nondescript pretty you know easy to miss but you click it and then it changes this toore posos excuse me and we need to click on verb right there second option verb and it puts a underscore and two lowercase V's let me zoom in again so we can see that very clearly there it is two v's those are actually V's in there're not it's not a w okay um good let's Zoom back out a bit and then let's go back to our instructions real quick frequency of LMA break as a verb now I'm going to um I'm going to click on the list here I'm going to keep the list as uh selected there and then I'll just click on find matching strings there and this is giving me um different forms of of break of the different word forms of the LMA break broke break broken breaking breaks and um it's giv me all these um frequencies by um decade across the ca from 1820s up to 2010s and here we have the all the the the sum of all those other sales to the right right so broke actually is more frequent than break itself about 3,000 by about 3,000 tokens there from 40,000 down to 37,000 etc etc etc cool that's how you do that okay let's go here number seven says the top five most frequent words words immediately following Mormon in now and their frequencies so now is a different Corpus so we're going to click on this um button right here on the top left of the screen that has a diagonal up Arrow it lists all the corpora in the English corp.org website and there is now now is news on the web okay it is a corpus of news on the web created by Mark Davies contains about almost 20 billion words of data okay and it runs every night it it runs and collects more news articles every night so it's it's like it's a monitor corporate it's growing okay so we need to find let's go back here the five most frequent words immediately following Mormon so we the word Mormon and then a word okay so the way you can do that you can capitalize this I don't think it really matters and then we have a space there so before when we saw gate we had asterisk no space and then gate and it found words that end in Gate including the word itself gate here when we have a space between a word and an asterisk it means it's going to treat that asterisk as a whole separate word okay so now if I click on I'll just click on the list here and let's just uh find matching strings and we find the results here and what are the five most frequent words we have church and then we have some punctuation here we have a comma period and a quotation uh double quote Mark now I need to explain that in most online corpora they're actually based they're really just a relational database on the back end and each little token like Mormon and comma and in church and faith and missionaries Etc are are they on their own uh row within a table within this relational tab database um if that's all Greek to you just understand that we don't have to worry about those few right there that have punctuation after Mormon so the five frequent words after Mor our church and Faith missionaries and Community okay in now good got that one done number eight the top five most frequent words that occur in between a and day in now and their frequencies okay so we just saw how to do that right we just saw right here that if you put a space between a word and an asterisk it'll treat that asterisk as a a word a token so let's reset our search form let's do a space asterisk uh uh excuse me day a space asteris space day and then we'll keep it as a list so we can see the actual results and then we'll click matching strings find matching strings what do we have here we have a good day a single day a great day a long day a bad day those are our top five U phrases with a word day and there are the frequencies right there in this column called frequencies there you go pretty easy right okay the top five most frequent nouns okay not any word but nouns preceded by uh or un and then an adjective okay we'll just stop right there we'll just look at this first okay the top uh five most frequent nouns preceded by uh or un and an adjective will'll just stay in now in the Corpus called now okay let's go back to our search form click on search tab I'm going click on reset to reset the search form uh so we need uh or un now that that little character right there between uh and un is called the pipe operator on my keyboard it's above the enter key I have to do shift and then the key above the enter key on my keyboard to get that character then I do a space and then I do an adjective again I can use a drop down let me just click on the drop down I can see that it said it has capital P oos there those are all capital letters so it's going to give me a full word so I click on that click on AJ and it puts in a space as well as capital adj and then a space as well there in the search form and then we need nouns again just go back here we're looking for nouns that occur in that position so again I could come back here and click on noun all and it puts in noun uppercase noun and we're going to keep it on list there and I'm going click find matching strings and there we go okay what do we see here a long time a little bit a long way a wide range a large number and their frequency is right over here on the below the frequency column good there you go there you go straightforward right I'm just going to point out something here real quick you can see the number of tokens it says total is 83.5 million tokens in the Corpus and it says unique those are word um types 56.9 if I round up 56,000 almost 57,000 there when where it says unique that those are word types these are tokens word tokens these are word types okay cool okay let's look at this next one here um and then the top five most frequent nouns that are not preceded by uh or un and then it does have an adjective before it so something other than uh or un adjective noun and now and the frequencies there's a couple ways you could do this you can search each one of those by by um by themselves with the minus sign so minus uh and then a space and then adjective space noun if I let that go you can and and get a result here in a few seconds and there are some results the other hand the other side in other words the federal government the American people etc etc and their frequencies right there um then you can put an a um a an in here UN in there and then search again click that button below it and then you get um these other results but you are getting some that have uh so just look for the ones that um are the most frequent that don't have a or on in there that's probably the easiest way to deal to do this the second half of this um number nine okay last one here says the top five most frequent sequences of adjective followed by noun followed by verb in now and the frequencies so adjective noun and then verb again we can uh go back to search click on the search tab we can click reset to re reset the form back to the default settings we can use the um drop down box that has a capital upper case there right P so I can click on I forgot what we're looking for we're looking for adjective noun verb got it adjective noun verb so let's go adjective bang noun boom verb bang click go adjective noun verb and we have current results set foreign language spoken Supreme Court ruled Human Rights Watch end of day data provided and their frequencies right there okay so there we have it that is um how to do these searches in with some of the corpora in English corpora thanks for watching catch you next time
Up Next

Advanced Searches in COCA: Academic Writing Tutorial
@dukewritessuite6673
710 views•2022-02-21

American English OKAY Over Time: A Diachronic Interactional Linguistic Study
@Abralin
1.2K views•2020-07-30

Forensic Linguistics: How Language Solves Crimes | PBS
@pbsstoried
1M views•2024-01-25

Accent Expert Explains U.S. Regional Dialects | Part 1
@WIRED
9.3M views•2021-01-21
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Linguistics

































