Diachronic corpora enable systematic, transparent, and reproducible study of language change by allowing researchers to track morphological, lexical, and syntactic changes across time through multiple independent annotation layers, thereby overcoming the limitations of traditional historical grammars that lacked frequency data, contextual analysis, and methodological transparency.
Anke Lüdeling: Diachronic Corpora & Language Change Analysis
Added:so welcome everybody to The Institute for digital research in Humanities occasional lecture series um when we have Stellar people coming to town we snatch them up treat them well and offer you a lecture that is digital Humanities inflected uh I'm Arian dwire co-director with Brian Rosenblum in the back of the Institute we and there Professor lling are sponsored by the libraries the College of liberal arts and sciences and the hall Center for humanities here at KU and we're really delighted to be able to co-host you today with um the MOX Kata Center just so you know there um is is yet another lecture in this Mar Marathon day um at 700 p.m. on scientific registers in German this talk however will be in English so please come if you have the time it's there at the moxa inst uh Center so very briefly um as you can see today we are having a methodologically oriented talk on language change using linguistic corpora uh I don't know if Professor lling is knows this but I've been watching her work for a long time I've even reviewed some things for you a long time ago um noce um and she is a professor of um Corpus Linguistics and morphology so linguistics um at the Hol University in Berlin she has uh an incredible record uh starting from uh as you probably know Germany is the place to be if you're doing Corpus Linguistics and um she's been all in all of those places where it is a place to be uh so she um is currently uh a professor um of German or Germanic linguistics she um is currently also the director of The Institute for German language and Linguistics institutes ofes B and linguistic and she is also overlapping a visiting Professor um at an esrc Center for Corpus approaches to social science at Lancaster University in the UK um and I have long been admirer of her work and I also uh happen to know that she's very very good at explaining things and so without further Ado I will let her do the explaining but let's give a very warm welcome thank you um and thank you all for coming here um Arian asked me to give a methodological talk but I believe that um you should always have a research question so we'll have two research questions and a methodological talk um before I will start with my research questions I want to start with two quotes this first quote is I I took one quote out of very many other quotes that I could have taken um it's about historical Linguistics and corpa and it says that the digital representation of um historical texts is just the same as we've always done before just digital there's no difference so we've changed materials from uh many times from Stone to paper to digital materials but it's just not anything else and in many ways this is probably true because many of the methods that we use in Corpus in for studying historical Corpus are the same that have always been used finding examples just we can just do it more systematically um we'll come back to this quote later at the very end of this talk and I will really start with this other quote that I really like it's by ham moel who's in uh the UK um he's a dietologist and he says data is ontologically different from the world this is really important the world is as it is but we can't study the world it's too complicated we have to make data out of the world data is an interpretation of it for the purpose of scientific study the weather is not the meteorologist data but measurements of such things as air temperature are a Text corpus is not a linguist sta measurements of such things as average sentence length are and I think that's really important so a corpus is to be interpreted on many levels and this is where a digital Corpus differs from all the other representations of data that we've seen before so if we look at moa's quote it says the data is um not the world but part a part of the world and already in choosing a historical Corpus a historical text or several historical texts we have made a a sample of what we want to study and in historical linguist Linguistics um very often the sample has been made for us by external influences we only have what survived we very often would like other samples very much we would like other samples but we can't have them but we just need to know that this whatever we have is just one sample of what we I really want to study we won't talk about that so much today but we'll speak about the second part the interpretation of the world um the analysis of the data if I come back to methods in historical Linguistics of course historical Linguistics has always been Corpus based because there is nothing else we only have a corpus of texts in a given language um and so we have pre electronic cor of course and they've been used for many linguistic questions such as I don't know gramar writing lexicography translation studies etc etc um and the methods usually are we look at the text we find an example we write down that example or an excerpt or something um classify that example and then use it for example to illustrate a statement of grma or some electrical enine or something um there's also been work on concordancing that is a word in its contexts in all these contexts in a given text and that's been done long before we had computers of course um it's mainly I mean we have we have concordances from the 13th century already and it's mainly been done for religious purposes so we want to find all the contexts in which the word Angel was used or something in the Bible or some other text that's important to us without having to read the whole text every time um and there has already for pre-electronic corpor we already find some quantitative research but not very much because as you can imagine it's hard how can you count you just have to read it and then when if you change your research question just a little bit you'll have to start reading the whole text again you you can't use what you had before it's really hard there's not much um I want to illustrate this by looking at uh grammars just a bit so we have grammars um traditional historical grammars of any language I will I will always illustrate what I'm going to say by using German examples because that's the only language I really know so I can only talk about that um and these but it's true for all the other languages that you can think of you can find analogous examples for English and and any other language so you typically have a language stage let's say middle hul do Middle High German and you have a grammar of middle H do and um it tells you a list of valid constructions and grammatical rules that were supposed be there in Middle do um it does not tell you anything about usage it typically does not tell you much about variation sometimes just a little bit it does not tell give you any frequency information it does not make it does not talk about the relationship between different constructions so if one constructions gets stronger over time that maybe um another competing construction might get um weaker and of course um history is written by the by the powerfold so we never get the loser constructions there might have been many but we don't know about them yeah and what's also important is um the way that many that these ramar writers who very often knew so much about their language stage how they arrived at their data and we see that in one example English translation on the next slide this is an example by one of my heroes actually although I will criticize this this quotation in a minute I want to say that hon PO is one of the best grammar writers ever but this is a very typical way of writing grammar so this is the English translation of what I just gave you in German um so hammont poell has written a uh at the the end of the 19th century and the beginning of the 20th century he written a five volume German grammar um and you can open it on any page and find comparable quotations I just opened it on any page and found this and it says sometimes the linking element as no matter what this is we don't care about I just want to speak about the evidence here sometimes the linking element s appears after feminine nouns although this is not yet common in written language and then he says compare the mind for so what do we see immediately it's a contradiction he says it's not in written language and then gives us examples from written language so what do we make of this but there are more more problems with this um first of all the world it's large he has a corpus and he doesn't tell us which slice of the world he actually uses yeah yeah so we know that all the texts he uses to illustrate everything he says about grammar a written authoritative text canonized text yeah um so we have novels by the great German novelists stuff like that we also find that there are 200 years between the birth I mean just for this example yeah 200 years between the birth of Life nits and the death of Hiner and another 50 years before Hammer so what is he trying to do is he writing a grammar of 250 years is that even possible he doesn't tell us he doesn't even tell us anything about this problem yeah so we don't know what this is we also don't know how he interpreted his data yeah he gives us a list of words to illustrate a statement that he made yeah um he doesn't tell us the definition of what a word is so here we have this genitive thing clusive which could be a word or not but he should tell us in some we have no context he doesn't explain the choice of these words to illustrate his statement um so we have no idea whether these are the only feminine nouns with an S Linker that he could find or whether these are just five examples out of 500 and that's what I said before he doesn't give us any kind of frequency information yeah um we also don't know in which context these Au is used be so he says they're not used in written registers so it might be it might just be that these authors use these examples in something that they considered spoken registers for example in I don't know their I don't know if they want to tell us any something about dialogue or something it's not the case here but this could have been the case and he should have told us so all these historical grammar look like this and you can immediately see that this is a problem um there is of course in historical Linguistics a lot a lot of uh progress since those times and we have have better theories about U languages and uh models of language and especially we have much better models of language change and one of the really interesting ones is variation ISM which is already um started by a social linguist Evol in the 1960s but it's is still a very useful model um and he said if we want to look at language change we need to look at variables and VAR variant and variables are abstract difficult to um Define this has been discussed in variation ISM since the beginning um and variant are concrete utterances what we really say and so if um H Paul wanted to make a statement about something that is maybe as he says already in spoken language but not in written language the variable could be uh certain feminine nouns in non-head position in compounds and the variant could be those nouns with an S or without an S and then we could see how these variants change over time in register and then we would have a real statement about what really happens in this langage so look at this this is I mean you've probably all seen this in your Linguistics classes so you have variables with v uh variant and the frequency changes over time and we don't know we have to find the influencing factors um and because we need to interpret the data for it to be useful we should really code every step of this interpretation in the data so that everybody can see it and that is usually not done I think it's a very very very simple and clear um wish I think it's actually a necessity because only if we do that our uh are our results reproducible or transparent and very often we can um interpret the same world the same data um in different ways and if we have a good Corpus architecture we can um write down the different different interpretations the conflicting interpretation of interpretations of our world um all of them into with regard to the same data then we can discuss our results exist um I think this is not possible so I I will speak this was my introduction I will now show you the data that I will use for my two case studies in which I will illustrate what I just said the data is a tiny dionic Corpus it's tiny it's by any standards tiny we made it as a proof of concept um it has four sub Cora I'll explain this to you in a second these are all religious texts because we wanted to keep the register as stable as possible which but as I said before the sample of the world is uh uh very often chosen for us especially in the older language stages we cannot find any kind of register in old high German because um there is only about a million text words of old high German has survived and much of that translations so we have problems here so um we have subcorpus old high German a subcorpus middle high German a subc corus early modern German sometimes also called early new high German and the subc corus German or new high German and um they're all religious texts and they are all manually annotated um with parts of speech um uh those parts of speech from theart Tuan TX set which is the standard TX set for modern German had to be conservatively amended because there there are parts of speech and and morphological categories in old high German like the Dual that just don't exist anymore so we had to add them um we also have sectional categories nemas and for ex and uh especially important for this talk we have trees that is a syntactic interpretation of the text they look like this um you don't have to understand everything that's going on here but just look at the P two words you see an old high German uh example and you have no that this is Middle High German by the way a middle height German example and it says and you see the modern German example that has g t so here you have an inversion of the word order and if you have a syntactic annotation like this you can find out this okay so I have two case studies now that use this data please keep in mind that this data is so data set is so small that we have to be be very very careful with any claims you make well what the first thing you always do when you have a new data set you explore it you have to get to know your data set um and that means qu qualitatively which you've already done when you annotated it but also quantitatively and what we've done here is we've used a technique that we um always use in learner Corpus studies um learner corpora are texts written by Learners of language for example Learners of German as a foreign language Nina constructs very interesting learner Cora but we've also constructed learner corpa and um in learner corpa you often have like the native speaker control Corpus and the learner Corpus and the idea is that the Learners of a foreign language should ideally behave like the native speakers so you get the frequencies for any kind of category from the native speaker Corpus and then you compare that to frequencies of those categories from the learner Corpus and you can find out overuse and underuse like they don't use enough definite articles compared to what the native speakes do whatever and we uh have uh a colorcoded scheme here and blue means underuse and red means overuse and of course you can um look at the numbers in the next table but it's much easier to see just look at the color ah and what I need to say is what we've done we've used this on our historical Corpus and here of course you don't have um like a norm but we've said we've decided that the modern German is like the end of the development up till now well we know yeah so we should U maybe make another study in 100 years but now we know the the result of the development and so we've used that as the Baseline and then compared frequencies from these other uh language stages so we get tables like this and you get them for every word for every syntactic uh syntactic node for every part of speech and sometimes they're just mixed up here the colors are not interesting but um this is part of speech from theart Tuan tag set and you can probably already see which ones might be interesting for language change studies and that would be you would always want the ones that have um a continuous um blue or red L I'm getting lighter so interesting could be infinitive main verbs or uh relativ visors yeah so you think ah they're highly underused in old high German A Little Less underused in Middle height German a bit underused in early modern German compared to Modern German so this could be an interesting development yeah this is just a diagnostic and so we look at we did this as an exploration step and then we looked at some of these categories in more detail so relativ visors are may be interesting um relative pronouns become more frequent over time that's what we just saw in that table so what does that tell us about the development of of relative clauses in German well structurally relative clauses are completely uninteresting because we already in every language stage we find relative clauses that look almost the same yeah and they also look almost the same as they look in English so I don't even have to tell you much about them so from the reward that he received from for his misdeed block yeah so this is a relative pronoun and they look in English they look exactly the same as in German shmid and others have claimed that relative clauses are the oldest dependent clauses in German and they also say that German has not changed with respect to relative clauses but we found a continuous development quantitative development again think back to the older grammar that would not give us this relative clauses have always been there always look the same so now what do we do now we will the first thing one always does is one checks the numbers yeah and the overuse underuse table I just showed you was normalized by number of token and the sentences we also know from other Explorations that sentences in old high German are shorter than sentences in mod German and maybe this is Maybe this might be a cause for this underuse so we also normalized our counts by Clauses and now we find it an interesting difference if we do that like retive Clauses per Clause yeah so we find no statistically significant difference between modern a middle height German and modern German they're almost the same but we find it highly highly uh significant difference between IDE German and so now our research question is different it's not like what's the continuous development but the research question now is what happened between old high German and later yeah and you can see that already using a different base for normalization the research question changes we now need to look into those relative clauses and um as you can see you always have a reference word and a relative visor you've always seen this before but if there's a change the change can be either in relativ visor maybe it can be in reference word or it can also be a change in function of something so that maybe now you can use relative clauses for more functions than you could in old high German because Schmid only talks about form so we'll have to explore these things um but the first thing is okay relativizer might be might actually be relevant because not only relative pronouns can introduce relative clauses there are other elements and maybe these other elements are those that introduce relative clauses mul so there are elements like the um um interrogative elements they wondered at what had happened to him save in English yeah so that's a relative pronoun and in Alli German we also find athetic that is not integrated relative clauses so you find this is really maybe um new to many of you um these constructions are not available in English or German now and said to them were there and said to them who were there we want something else yeah that the the introductory word to the relative Clause is just missing now um the thing is we have a syntactic annotation so we don't have to rely on the form of the introductory word to a relative Clause we can just count relative clauses I could have done that from the beginning because here in my syntactic trees it says ver fls I can count those and if I do that and then I count the other variables relative visor I can see a difference and don't be fooled by the Numbers these are the uh absolute numbers but if you normalize that again by the size of the corra you still see that um the Rel that the whole concept relative Clause with any form of relativizer is highly underused significantly underused in all so that didn't answer our question but we did have to explore this then we need to look at the function and we also need to look at the references the function the we look at functions M the main functions of relative clauses are um restrictive and a positive the same as in English and they are already there in Al German so the functions haven't extended doesn't help us either no answer to our resarch question then what the answer is so simple all German has no nouns almost and so we have fewer chances for relative closes of course Rel Clauses can modify other things but they hardly ever do even in modern German relative clauses most of the time 90% of the time modify nouns and old high German has fewer nouns and if you um normalize everything um by nouns the difference goes away and is white no change and why did I tell you this this is a methodological talk as requested by Aran and we we actually really went through all these steps to arrive at this conclusion it took us a long time um our initial suspicion given by the overuse underuse table seemed to be a good one huh there's a continuous change look at this it's a nice diagnostic for change and then we triy out tried many many different little sub isues and we couldn't find any interesting difference and then we saw that the difference that we found was due to a completely different difference and then of course now we can explore why are there so few nouns in old high German but that's a different question and one that I will not talk about later we not talk about today though next case study the paraphrastic perfect modern German has several tense forms that can be used to refer refer to past events the forms look like they look in English the functions are different so don't be fooled by the forms so we have a predate I worked like the past tense in English we have perfect form and the perfect form can have two different kinds of auxiliaries H and Z which are the same for our purposes here um in the early stages of um old P German we only find right um slowly the analytic form develops we can have there are different theories about why that happened which we don't have to talk about today um both forms are still there today both forms compete in some readings and that's where the difference is to English because you in English there is a clear um usage difference or semantic difference they don't compete in German there are many contexts where you can really use either one and it's not a grammatical difference so now look at the numbers again again from our tiny Corpus we in our Corpus we didn't find any perfect forms in the old high German Corpus but look at this in Middle High German we find some perfect forms here and in modern uh early modern German we find mostly perfect forms so my question that I cannot answer if I look at this I mean that's the question everybody has immediately is why do we still have the predate it should have been gone the next day I mean look at this yeah it should be a loser construction it's still there I don't know why but what happened between Middle High German and very mod in German something dramatic must have happened with respect to 10 fors it's only dramatic for I guess I mean have this kind of change yeah um in the literature sometimes claimed that the choice of or the the the choice the early choice of um uh perfect over predate is triggered by aspect again different from English this might be a good explanation for English uh past tense perfect tense division but it has been claimed for German in a different way yeah it has been claimed that the in the resultative we use the perfect and in the non resultative context we use the pr it's also been claimed that it's been um triggered by formality so an extra linguistic influence maybe in informal colloquial spoken context we find it perfect and in written and formal context we find a preent you still find these statements in modern grammars and they're still wrong we found this in a purely written corus a dramatic difference so this statement as it stands cannot be true but um here my main methodological point in the second case study is that sometimes in the course of a study you need to add new annotation layers because now we need to find out whether aspect or formality play a role in the choice of perfect inate when we made a corpus we didn't know that we would ever need this so we want a corpus architecture that is flexible enough so that we can add this annotation layer because we now need it and you can immediately see that or just imagine you had to um go through a text and decide this is a formal part this is a less formal part this is less formal part this is less formal part this is again formal and you can see that if your neighbor does it also you will not agree everywhere it's a difficult interpretation this is what I mean by interpretation and typically the things like these have been done away from the Corpus on a piece of paper on a fire card in an Excel sheet in a word file whatever um and people publish the the results but we don't know how they distinguished formal from informal compex they don't give us the full story but you can immediately we see that this is is going to be controversial yeah same for aspect what are what are perfect what are resultative contexts and what are non- resultative context sometimes it's clear sometimes not so we had to annotate the Corpus for formality and for aspect and it is there you can now see in the Corpus it's publicly available the whole Corpus is completely freely available how we did it and you can assess whether you think we did it right or not okay so we've divided the context into resultative and non- resultative every sentence and also into narrative and communicative they're all written they're all the same text but there are sermons and in those sermons sometimes the preacher um tells us about the life of Jesus we call those narrative and sometimes he speaks directly to the audience and say you you should not sin so much and be a better blah blah blah so he speaks to them we call those contexts um communicative because we had to find some kind of Distinction because formality cannot be it because it's the same text and the same audience and the same speaker I'll explain how to read these figures these are figures that give us three three uh types of information at the same time and it's why they look a bit more complicated um so we have Middle High German and mod early modern German and then we have percentage of perfect the fact percentage of perfect use of all tenses um the black um bubbles are narrative cont context and the gray bubbles are communative contexts and the size of the bubble tells us how many of those contexts there were that's why we need the size because otherwise what we find in early modern German that in narrative contexts no in Middle High German um we find in narrative context almost exclusively the use of the parid a very low fraction of perect and in uh communative context we find about half of the time you find perfect so this might actually confirm what was said about formality maybe if communicative narrative has anything to do with it yeah but it's already over here yeah so what we see is the perfect takes over all context in 100% of all communicative context we have perfect and almost all of the narrative contexts perfect so this explains some but not everything this is what I just said um but this Factor alone does of course not explain this huge dramatic change so we did the same for these other context resultative and non- restive context and now you already know how to read these uh figures and um you can see that in Middle German in all of the non- restive contexts we find find the perate and in some of the non-res resultative context we find the perfect again change seems to involve both contexts so this explains a bit more so we had we had we saw that the communicativeness of a context explains part a large part of that um dramatic rise of perfect so perfect takes over all those contexts it couldn't be used before in those contexts it can be used now and another part of that those contexts and another part of that change is explained by um resultative and non resultative cont so so together um most of that Chang um what's not explained and what I can really cannot explain is why we still have the pred it should not be there have it but why it where does it survive I mean this is okay I I really have no answer but given these things given the fact that the perfect takes over everything Why didn't it take over everything okay so my summary here um and uh my methodological point is the important point you can forget about the predate and the perfect immediately but what you need to remember is in the course of an investigation it might be necessary to add more annotation layers and because these can be highly uh controversial you need to add them in the Corpus don't do it somewhere else we need to see that that's that's the really important point and that's what's usually not done why how can can we do this and now we get five technical slides that I just need to give you because we need to find out how we can do this so we want many layers of annotation and these layers of annotation that is interpretation of the world no um can be in different formats there can be token annotation that means a token is just because everybody know what a token is that's just a graphemic word or something similar to a graphemic word or uh a string between um um spaces or something if you wanted that we have span annotation that spans several tokens we have directed Asic graphs that's for example Ina trees and we have something like pointing relations that is anaphoric change chains for example and we want to be able to add new layers in the course of an investigation whenever we think they're necessary so we don't want a corpus that is fixed we want a corpus that's flexible so the vision is here yeah the vision says we want to build an an and store historical and dionic cor dra in a way so that they can be found and analyzed by as many people as possible for as many research questions as possible because it is difficult and expensive and time consuming to digitize uh a historical text this involves clear Corpus design depending on the research question well despe metadata that you've probably all talked about many times reusable in standardized formats um this involves a multi-layer architecture um where we can add annotation in all the relevant formats that is token span trees ET but um okay but I'm not saying what you should add I'm not saying what types of information you should add I'm just only talking about the formats we also want in the end I'm not talking about that now but tonight the representation of intertextuality um we also want reliable and transparent depositories and that's the hardest part and the part unsolved at least in Europe but I think also here nobody will pay for it because it's not just storing it's also curation Etc we want a conversion tool that converts a format to many other different formats we want multi-layer search tools powerful multi-layer search tools um and we want as much computational help as possible and um in Berlin now I have two uh two slides about tools that we built in Berlin but they're all freely available so you can use them they might be interesting to some of you um the first um product I would speak about the repository is the repository for historical cor it's already there you can look at it um you can display cor Pro so if you have a corpus a small or large Corpus you can just put it there and then people can see it and find it and that's the first thing we need to find it we know what you have um you can search them the ones that are there um you can download them and add new annotation that is really important that's what I just said I mean sometimes you want to add annotation to a corpus that you didn't build and then you can upload new annotation or any other new corus or version or something um the cor can be in any format and they can have any number of annotation layers as many as you want none for example can still be very useful to have just have the text yeah um and the only or only point where we standardize is the metadata of course so that we can find um this is a typical computational picture I explain it um what you then need I said you can store any corpus with any format and then you can download it and add annotation well but we know that each type of annotation or there are dedicated annotation tools for each type of annotation and there are almost no tools that can do different types of annotation there are some that are being developed at the moment but not yet really good enough so you take one and you want to add token annotation and you take the same Corpus at the same time without her knowing and you want to add syntic anation and both of your tools will output different formats and so you need a conversion tool that can then put this together again because you don't want to deal with this because you didn't even know that she also works on this cus yeah and we have a tool that's called salt and pepper and that can do exactly that so you as a user don't have to worry about all this you can just use your dedicated annotation tool that you always used and wanted to use add your annotation layer store the Corpus back into the repository and then salt and pepper will put that into the Corpus add that to the other layers the other layers are technically independent of each other so nobody knows and then you can find it in the lau repository or wherever but you can also search to all these annotation layers and that's done with using the search to Anice that we also built um Anice is uh a very powerful an uh Search tool for uh deeply annotated War it takes a minute to learn and that's because it's so powerful and general because we don't know which corar will be there we can have corar in all kinds of interesting languages yeah left and right script right to left s whatever you want will you be good um and you can search to all the Met metadata so so you could then also use her syntactic annotation to search I don't know for any something that's important to your research p and that's the way okay so s so we started with a quote that said and a digital computer is only the latest in a series of tools blah blah blah and we've heard this before so we can do what we've done before maybe faster maybe more systematic but I think that's just wrong I think we can do completely different things now because we can make our interpretation of the data visible for the first time we can publish our interpretation and for the first time we can make research in historical Linguistics like Linguistics um transparent and reproducible it was never before possible so I think that's uh something completely different from what we had before okay so what I said is that we need to categorize and interpret our data categorization is interpretation and every quantitative analysis depends on a previous qualitative analysis what I said is a desum we wanted to code every step of the analysis directly in the data and I also believe that annotation is research while you do The annotation you understand your phenomenon that's a different talk though yeah um and okay and then for language change studies I want annotation layers that are variables with variance um I want to think about the unit of normalization quantive normalization um and I want to look at context factors of course um yeah that's I said that before these two points and I'll stop with the MRA that I always use and it says Corpus design and annotation depends on specific research question but Corpus format and architecture follow general principles thank [Music] you please questions comments made an interesting pass comment on the issue nobody wants to pay for the curation of of data but yet we're talking a lot about open data and and there are National corporate and so forth so I wonder if you could maybe say a few more words about how you that cont how it's developing in European maybe okay I I I I know a lot more about the European situation than about the American situation so maybe I should not talk about the American situation um there are uh several issues that are important um one issue is plain storage I mean you can store whatever you have on a hard drive or something and many libraries now offer a service of storing things and then trying to store them on New Media when the stored media become faulty or something but that's only part of the story because um you also need somebody to reprogram these programs because the formats will become outdated and we have no money for that second step the first step is probably not the the difficult one and so in Europe we have a number of European um initiatives like clar and Daria I don't know about whether you know about those Aran knows that they're uh paid for by the European Union and they Dar claring is for Linguistics and Daria is for all the other Humanities I guess I mean that's a simplification and they they are supposed to explore exactly this issue what can we do how can we store data how can we curate data how can we keep it Al um and they have no Solutions and they are funded like almost everything else they're funded for 5 years and much of the research that I'm doing and the resour that we are building are also paid for by project money which is sometimes 3 years sometimes 5 years sometimes if you're lucky 10 years maybe but it's never um um sustainable it's there I have no I have nobody to give me money for the next 70 years to do this and and we don't know how the format will develop so I think we can do several things to to uh make the situation a bit less problematic one of the things is a converter framework that the one that we have so that we at any point we can store we can convert the data to other formats so if one of the formats maybe um falling out of use there are others that might survive longer but then somebody needs to update that conversion framework because we don't know what format will be the format next year so that's really a problem also pro programs like Anice and other search um tools need to be updated it's not not only a question of storing them somewhere and so we have this problem and as far as I know uh American sitation is not [Music] better I hope I'm wrong I hope that somebody finds a solution with this because that's the the real other questions with the with the slot corpor we things are really complicated because we have National corpora and they make it hard to extract data actually yes the Russian national Corpus will not give you all of your hits it will only give you um you know a sample of them and the check National Corpus also doesn't exactly make it easy to to get stuff out in a usable way and often they don't give you enough context yeah um I was I mean I find myself having to do what you say we shouldn't do which is put basically come up with a bunch of predictor variables in an Excel spreadsheet because I don't have I don't have corpora that have what I need so I take text and I go through them and and um I think I don't do you know about the trolling repository up in the University of Trum So they come up with what you're doing for corpora but for databases so if I want to put my database up there so people can see what my particular variables are and it's I think English and and German are have it easy no but in German the national Corporal just like you describing yes and it's a shame it's terrible and that's I I I could go on and on about this at least you have like sermons sermons are great for various like communicative and narrative but but we had to make this Corpus oh you oh you had to make it and we made large corpor too and because you do you have any recommendations for OCR and constructing the initial yes I do because I don't have time to do to do it at least the ways I don't have to do it so maybe you can tell me afterwards or something how to make it easy yeah it's true and I know that for for Slavic languages some of them Olan Mayer for example in Berlin he he construct corpor that will be like this that will be in the loud repository and that will be freely available with that I mean if everybody were doing it I wouldn't have to talk about it Soo much but the fact that what you're describing for the Russian national Corpus is a situation almost everywhere you typically you cannot you can you can only get one sentence you cannot get the cont context very often you can only look at things on a web interface but you cannot have the Corpus to annotate it further yeah and and all of these corpor are publicly funded and it's a crime yeah yes because you cannot do your research in a transparent way and in a way that's that's up to standards and you it's not your fault because you can't have the data and some people can do like secret research maybe and it's the same in Germany and um that's why in my in my uh Department we are building Cora and all of them are freely available in any of these formats because I don't believe in secret research but yes and we all need to talk about this all day to everybody and tell them that they can't do it anymore and yes I can tell you something about OCR and ID can also help with OC arms but please other questions comments Curiosities anything you wanted to know about Corpus Linguistics but we're afraid to ask I have one question about hisorical corpor the texts you've chosen have a pretty clear like historical provenance in sort of you know the year pretty precisely what about working with text where maybe there's a century r how you determine whether to work from the beginning data that the end dat that orwh mhm um I think that depends on your research question actually that's what I said so Corpus design um and what you want to annotate really only depend on your research question that's no General answer to that but you need to to just put that into the metadata and there are metadata schemes like TI or so that really allow you very detailed um detail information with regard to everything and um um if you ever find the time to read through the ti manuscript part has anybody ever done so David probably he made it didn't he but other people Michael yeah maybe but um but yes you can represent that but how how how you will interpret that that depends on your research question so I have no General answer to that so I don't know what you're trying to find out but you should represent what you know and you should represent the range and that is certainly possible the other comment I'd have is ID puts on workshops in September and March and we are suggestion based and so if there is an interest in learning how to build a corpus or assembling metadata for a corpus or marking up texts for literary linguistic historical Etc purposes um please see Brian or me or shoot us an email if there's at least two people that means there's probably more people out there and we can put our Workshop I guess it's about corporation that don't exist for them and so often times creating like these materials as you go um and you mentioned that even with this any you showed us that you wanted to control certain things like register you have you know type of recommendations for what type of things you might want to control for if you creating around that M so is that modern language or historical language modage um so we in in in our group we' have done not not me personally but and other people in in the group have collected for example the house Corpus um and that might be a similar situation maybe to the one that you I don't know what languages you working on but similar like that um the thing is very often you don't have the choice depending on the language you you often have I I either religious text or chat and depending on what registers that language is used used in I wouldn't control for it I would take whatever I get we call that opportunistic collection but I would put everything in the metadata so that then in a in a tool like Anice um but many tools can do that you can have ad hoc sub so take everything you can get and then concentrate on the metadata and then build your subora according to your search question because you never know beforehand what you like want and um I mean of course nobody can write a grammar that deals with the Bible in chat but maybe some parts of your grammar might PR takeable well then thank you very much
Up Next

Using Large Corpora to Analyze Language Variation and Change
@laelwebinars9025
628 views•2020-06-20

American English OKAY Over Time: A Diachronic Interactional Linguistic Study
@Abralin
1.2K views•2020-07-30

Forensic Linguistics: How Language Solves Crimes | PBS
@pbsstoried
1M views•2024-01-25

Accent Expert Explains U.S. Regional Dialects | Part 1
@WIRED
9.3M views•2021-01-21
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Linguistics




![Ferdinand de Saussure • Curso de Linguística Geral [Primeira Parte: Princípios Gerais]](https://i.ytimg.com/vi/HLZr4de2XNY/maxresdefault.jpg)





![[2025 Spring] Natural Language Processing for All: Introduction to NLP with SpaCy](https://i.ytimg.com/vi/BTak610q6H0/maxresdefault.jpg)



![Penelusuran korpus beranotasi POS tags [LK 107]](https://i.ytimg.com/vi/Y9wTSX_KWas/maxresdefault.jpg)



![Mudança Linguística - [ Unidade I - Aula 03]](https://i.ytimg.com/vi/V9ih_2f9z3k/maxresdefault.jpg)




















