Effective language documentation requires understanding three interconnected components: metadata (data about data) that provides essential context such as who, when, where, and what language is in recordings; proper file management through semantic naming conventions (including language code, location, participant, date, and content) and organized folder structures; and annotation tools like ELAN for adding transcriptions, translations, and morphological breakdowns to audio/video recordings. These practices ensure community access, long-term preservation, and usability of language data for both academic research and community revitalization efforts.
Language Documentation: Metadata, File Management, and ELAN Basics
Added:good evening everybody hi folks how we all doing let me just get the Facebook live stream started and as always if you'd like to tell us where you're watching from say hello to this week's presenter please do get in this Facebook live stream started so that people can watch over there if they want to here we go and it is live all right hi from Ethiopia thanks for joining at this late late hour all right for those of you who are just joining feel free to say hi tell us where you are joining from I see some familiar folks in here tonight all right cool well it is a little bit past the hour hi from New York and Iowa and Westminster BC Indonesia excellent hi from India Algeria gosh yep we got an amazing range of places again this evening morning afternoon hi from Sudan cool well anajima you've got a very International audience and a lot to say so without further Ado I'm delighted to hand it over to anaji masakia who will follow up tell us a little bit about taking good care of our language recordings welcome onojima thank you Anna um of my audible am I clear yep you sound great okay perfect so this is uh it's early morning here very early morning in fact so namaskar who prabhat uh that is the asmies for um hello good morning I'm anujima saikya and today I will be talking about metadata file management and the basics of anon so this is supposed to be a very it's it's supposed to be an introductory uh Workshop into how we conceive uh these three topics or yeah how we basically conceive these three topics so you must be uh shocked not shocked a little baffle uh more or less the same as to what is a dog doing here so this dog here is almost like my Alter Ego uh so uh this talk is exactly my face when I was in my masters and when I was actually told or when I realized that field work is not only about meeting amazing people uh living with amazing people experiencing their lives uh learning a new language Etc but a field work actually entails learning softwares doing those threaded Excel sheets and transcribing and so I was like really is this what I signed up for okay so we go to the next slide and I'm still equally confused do we really need to learn all of it and uh after so many years of doing field work and learning more as the process continues as I learned there was a very important point that I took home and that is that a resounding yes a definite yes we need to know what is metadata we need to know how the how transcribing softwares work and the earlier we know it the better it is because as I already mentioned here it is one of the few ways actually and it might come shocking to a lot of people but it is actually one of the few ways in which we can truly contribute to the community so um yeah so what do we mean by that let's get into uh why do we think that documentation matters so here uh I'm speaking about the Penobscot situation and I'll start with this to give you a brief introduction as to why does documentation matter so pinobscot is an indigenous American language of the Eastern branch of the Algonquin language family was traditionally spoken in uh mine but currently it has no first language users so um like many other Native American communities in North India sorry like many other Native American communities uh starting around the 1800s what really happened was um uh there was a criminalization and a punishment which was imposed on students um who were speaking their ancestral their ancestral language so they were living in government boarding schools and by the time they return back to their families they had completely forgotten their mother tongue they had completely forgotten the language so in the past 40 years in the pillar in the Penobscot Community within the uh pinup sorry within the Penobscot community members What was seen that there was an increasing number of suicide and um uh somebody working on uh the language revitalization of the Penobscot Community said I quote is that we internalize oppression as a result of which we have high suicide rates and so this was the situation was which was ongoing there were no speakers of the language remaining um there was no shared history to which uh the people in the community could go back to they were speaking a foreign tongue because of which the the internalization of Oppression yours of Oppression was of course traumatic which need led to an increasing suicide rate so what Carol Donna did um who I just quoted she uh she compiled an anthology called these still remember me which has 13 traditional tales about the tribes cultural hero who was known as kluska so here on your left you actually have an image of glisca so uh this person here is close cup and you can see him converting another person into a tree and this as I think you can see here sorry my frame is a little too higher uh yeah just a second yeah yeah as you can see here it is sorry it is um it was on the book it was on the book but sorry I'm so sorry I'm getting the bar underneath yeah it was a scrapping on the book back by Thomas Joseph of 1884 and the way I got it was definitely through one of the archiving sources so what happened was uh because um in 1918 there were certain Penobscot Legends which were documented as uh in terms of phonetic transcriptions by an anthropologist called Frank Speck and that was uh archived using those documentations from 1918 there was a dictionary of Penobscot which was produced which is in line which is being produced currently so what we saw is something that was done in 1918 so the situation here till now is that uh for the longest time we had a language with no uh speakers it was only the Elder generation then the then uh the generation before Dana nobody was speaking the language and by the time it came to Dana there were just a couple of uh people of her grandmother's generation who was speaking it but because in the 1918 we had Frank's pack who had started working on this language we had some archival sources and because of which in the 2020s 2022 even currently there is a linguist called Connor Quinn who's actually working on the dictionary of Penobscot so coming together of both of it the dictionary plus um the the myths and the folk folk tales of gluska together brought in this idea of a shared identity and further pinop Scott is now being used as an act of decolonizing and healing so in a society which led to such high suicide rates Now language and uh the folk tales are being used to bring in um an idea of shared identity and it is very interesting because I was reading as you can see here um I did provide a link for a new yorker's um article and here actually Carol Donna who's been working with the revitalization work of Penobscot said that um it's amazing how in Penobscot you can actually speak full phrases and you can uh narrate Expressions through just single words and she gave us some examples as well for example uh lunch is noon eat butter is milk crease canoe is that which flows lightly on water so after speaking English for so long when Penobscot speakers went back to speaking it or reading stories which spoke about these words it brought a moment of joy for them and this is how uh the revitalization situation in Penobscot has been looking extremely um empowering and it is being used as a tool for decolonization another example of an Anglo sorry or another example of um and uh Algonquin language is the uh one panop people and the one panag language so back in 1903 there was a dictionary which was uh written and that dictionary was modulus derived out of the iliads Bible which because of certain archival sources because there were um digital copies of it in 1993 um um a woman called Jessie Little Doe pain she actually reconstructed the pronunciation grammar and vocabulary uh to to to um to sort of visualize uh the wopanak language Reclamation project and it has been one of the most um uh it it has been one of the most successful documentation or revitalization of a language which was quote unquote dead so uh in 2014 a handful of children were actually growing up as the native speaker of the wampanook language so these are two very important examples onto what could happen right if you have uh archival sources here we saw uh bits of archival sources which were sort of um reconstructed to form entire grammars entire vocabulary Etc but then there are also situations where this isn't always the case um I'll go back to the next slide and this is from the state I am from which is Assam in the northeastern part of India and here you see an example of the Thai manuscript that I a home manuscript so Kaya home has a similar story like uh wampano where in um because it was the language of the Kings they had the manuscripts and the manuscripts have been preserved because of which even though Taya home is now considered dead there has been a lot of work now on reclaiming Taya home but uh for instance another Thai language which is Thai Nora So within Thai languages of Northeast there are seven uh varieties or seven languages out of which uh three of them are have a lot of speakers they are um spoken uh they are spoken there is a lot of intergenerational transmission uh but um there are a couple of them which are slowly uh getting extremely endangered in fact I come young one of the languages out of the seven in the group has less than the last time I went there it had less than 35 speakers so um the situation of Thai Nora which has which could be included or excluded from this group is a little critical because the only documentary proof of the existence of tinora is by uh grierson Back In 1902 where he did mention dinora but we do not know anything more about it uh at that point of time in fact in 1902 as I've mentioned there were only 300 speakers so now there are propositions that probably tainora uh or speculations that probably Tai Nora was the same as time kamian but we do not have conclusive evidence or enough archival data to actually prove or even disprove that so this is all speculative so there are a lot of languages across the world where due to a lack of archiving uh or a lack of code archiving data we in we could lose them pretty much in the future um okay so let's go back to the let's go to the next slide we at this point of time uh we want to take a pause for a bit and look at these two sentences here number one is what is data for us is the lift experience of a language speaker it is an important index of their identity number two is when we enter a community we disrupt their everydayness so what do we mean by these two uh sentences um as linguists or as Outsiders entering into a community um we look at somebody's culture somebody's language somebody's identity as a cumulative whole called data right data the even the terminology is a reduction of somebody's lived experiences and it is an in fact an important index of their identity so as an outsider when we enter into a community We There is a general flow of things of life and we disrupt that flow we disrupt that everydayness and this should be or this should ideally be the starting point and the starting point of a researcher entering a community wherein they are an outsider is that of a disadvantage and this should be a very very important point and this should be a reflective point uh so we start at a deficit here as I said and it is our responsibility to prioritize the community and give back archiving of data hence is one of the most crucial ways or a crucial step in actually trying to give back to the members of the community and it completely depends on what the members really want what kind of archiving they want how do they want their data all of that even though we'll come back to we'll come to it in a bit but all of it are very important considerations that we need to take a moment and think about right now before we get into the nitty-gritties of terminologies data in hard drive would die with us or it would just stay with a couple of people who we can the hard drive to so this is where exactly and the reason why we need to follow this step we have data and our process should be to Archive it data is of two types to give a very very large classification or diversification primary data and secondary data I'll get to it in the next slide but for now it's divided into primary data and secondary data um we at all costs need to prioritize Community access to the data say once we collect it and uh it is very important that we speak to the community members see the stakeholders which necessarily not be the elders but most time it's the elders The Gatekeepers say the village headmans pradhan's Etc uh and the members of the community in general we sit down with them and we ask them categorically as to how would they like their data to be so we need to prioritize Community access to the data we need to find a digital archive or a repository that will accept uh your documentation data or you can also make your own arrangement with that archive so a lot of universities have um uh repositories or they have um tie ups with archives Etc um so that is something which has to be uh found out before you go into your field work or else if none of that exists you can make your own arrangement with the archive you can talk to them uh you could see um you could you could see uh how are the policies of the archives so on and so forth uh going to the next slide so here is a list of around 60 Archives of the world I just listed it down to uh for everyone to more or less have an idea of uh where they are how they are if you want I can also come back to it share it again and uh different archives also deal with different parts of the world for instance me being in uh uh in in South Asia there's course all that works a lot um then there is also the endangered languages archive which works with uh data from all across the world and so on and so forth um going back to what really is data a data is a collection of numerical or textual observations which can be clean processed used and analyzed for many purposes as I'd already shown before there's primary data and secondary data so primary data is the raw unedited and transcribed uh version of it um it is the very raw audio or the video or the original observations that one has the secondary data on the other hand includes the transcriptions translations morpheme breaks annotations so on and so forth okay so here we have an example of a data for instance we say that this is uh uh an example of data from the world of Art so what do we see here there are two uh major observations that come out when we look at this the first is that we see a man kissing a woman and then we see some rectangles and flowers and those are the two basic things which comes out also of course one can say other things like there are little concentric circles in between uh there is also a lot of yellow and the color patterns and so on and so forth but at the start of it these are some of the observations that come out uh but suppose we go to the living room of somebody at some a friend's house and we saw this painting do we really understand what's the background of this painting or what is really happening in the space I don't think we understand what's really happening so this is where we bring in a little bit of description through the very basic of a descriptive metadata so a descriptive metadata here then gives us a little bit of detail for instance the title the title of this very famous rather very famous painting I'm sure a lot of you all know about this is called the kiss we see the Creator here being mentioned who's Gustav Klimt the d8 created 1908 1909 the physical dimensions the type and then we have the mention of Belvedere here which probably looks like the um the gallery where it is stored or the archive where it is joint so this is the basic descriptive metadata where we kind of know what is really happening a little bit of detailing now let's go into this here if we suppose take this out of our drawing room or a friend's drawing room and we want to locate it in a gallery in the gallery we have to then associate it with other paintings so how do we then do it first we could probably put it with the paintings from the 20th century because as we already saw before it was around the 1908 or 1909 uh we could also put it as under symbolism which was a movement a big movement an intellectual movement a movement but then uh art and culture that had come out at that point of time there were other artists who were part of this movement there was Frida Kahlo there was Pablo Picasso there was Albert Munch so probably if we put this painting of Gustav Klimt with Frida Kahlo or Pablo Picasso's painting maybe it will make sense the third was the Vienna secession which was um another movement so probably if there are other painters from that time for instance Otto Wagner or Joseph Hoffman we could also be put the painting with them so this is how metadata really helps so we have the data here which by itself doesn't convey a lot but when we see it with a little bit of description the descriptive metadata or when we see it in association with everything else around happening uh during that time we can contextualize it so this is what we call a metadata ideally research data should be accompanied by a supporting document also known as metadata that helps to explain what the data are about so in a very stereotypical uh description it's known as data about data um so that is broadly the idea of metadata in language archiving we use the term data very broadly to include any recorded observations of language including sounds images writings that can be transcribed translated lost watch read listen to or measured for analysis okay so going to the next slide I would want to just give me a second uh I would want to give a metadata sample here let me check if I can okay so can everybody see the sample uh and out you see the sample here look there it is okay so this is in fact actually um a metadata example from uh kaipolyohone which is the archive of the University of Hawaii at Manoa so most archive would have their own metadata sheet uh with a pre-generated format for instance this might look a lot right in fact it's it's even um okay it's even a lot more you know um just to give you a brief example see it just keeps going on and on and on so at the first glance it could be a little intimidating but um to be honest while you build through it it's not actually that difficult so what really is happening here is uh here you have the item name right this is uh Anna's abbreviation and the number here so here you see it's all in one sequence right there is a complete numbering and the reason behind this is because all of it is actually a video file all of it is a wave file that you will see uh probably here yeah yeah you see digital weight pipe so because all of it is just one kind of recording it makes sense that it's all listed according to numbers we will definitely come back to um how the file should be named what are the conventions Etc but for now to introduce some metadata this is how it broadly looks you uh have a file naming um convention here then you uh write the con the name of the language the language code if there is one if there is a ISO code you put it here um uh then you write the the name of the languages and again uh the iso code and so on and so forth there are also um a description of the area the description of what's the interview is about who the speakers are who's the contributor and you can go on and on about and you can create a lot of columns here depending on uh first the archive for instance if they already have a metadata list good you can just keep on filling it accordingly if not you can also make one from scratch which would which would also obviously be a lot more time consuming but that would mean that you can also put in your priorities depending on your naming conventions okay so this is the very basic of metadata if you want me to come back to it and explain some more by after uh the end of the talk I could always do that but going back now to the uh presentation so that was the metadata sample so going back to our data now we would broadly look at how uh data should be um should be sort of organized so that it's easier to be later on uploaded to an archive so we see here that there is something called a session so uh generally what happens what precedes a session which I haven't mentioned here are resources so resources are part of a documentation uh let me give you an example for instance um there is a session of recording going on where uh say you are describing um um say the folktale a folktale uh in a particular language right so while you're doing that during that session there would be say a video recording an audio recording there'd be somebody clicking photos so on and so forth so those individually a photograph um of of that of the recording see or um the audio recording file and the video recording file would be in uh would be individually known as resources and these three resources would come together to form a session and uh a corpora is basically a specialized it's a type of a collection actually wherein it is largely used to build uh dictionaries which is it is mostly it has a it has a purpose that is why a copper has been and the most important thing here actually is the collection because uh the session the copra all of it forms a collection and the collection further goes on to form the assembly ages so I'll start from the assembly just so why what is the assemblages the assemblages are the combination of everything you have collected Let It Be photos Let It Be videos Let It Be audio whatever all the data that has been created are assembly ages but then when we curate the assembly ages we call it a collection and a collection is exactly what goes to an archive not the assemblages because suppose you have clicked 10 pictures of the same recording event or the same event right in one pick up one picture the first picture is at a 45 degree angle the second is a little blurred the third is say somebody comes in between or whatever so on and so forth out of those 10 images you would want one image to go into uh the collection the one which you think comes out best out of it but in the assemblazers all of it go so why do we need to curate the assembly ages to go to the collection that is primarily because archives uh do not have that much space as well that you put all everything that you have collected so we need to carefully curate it make the collection for it to be ready to enter the archive so just to give you a brief breakdown this is what a collection is we'll come back to the collection in a bit again uh now we will quickly go to file naming and file Arrangement which is very very important because this one step or other two steps would actually uh actually save a lot of your time a lot of your effort and a lot of your frustration in the future we just went through the metadata sheet right and what we saw here oh sorry what we saw there was that there was sequential ordering because it was all the movie uh the wave video sorry the media files which were wave videos so there was zero one zero two zero three that was sequential ordering but suppose we have an example here right we have cooking Story one cooking story two cooking Story three so what do you think is the issue here for me I think the biggest issue with naming something within a sequence is that it doesn't tell me enough when I put it up in the archive and say somebody 100 years down the line looks up for it they would be so confused they'll be like okay uh what is this cooking story about is this uh where in what community was this cooking story recorded um when was this recording was it recorded a thousand years back what is it recorded 200 years back when was it actually recorded so there's absolutely no more information so for sequential ordering it actually makes sense when you're sharing it amidst uh your colleagues co-workers then the sequential ordering makes more sense but it is actually the semantic ordering which gives us a lot more insights into the organization of data so how uh semantic ordering comes in is probably a little like this we have the language code we have the location we have the participant name we have the date and we have a little bit about what it contains and then whatever the extension is for here it's wave so I've continued with wave because uh cyrium the language that I'm working on doesn't have a language code it doesn't have a ISO code to be particular I have written it as cyrium I've written the phone the location where I've collected it is balisor the participant name from whom I collected it is faka eye Mall here is the date so this is generally the date convention it is year month date and this is very important to follow it uh then it's cooking Story one dot wa so at starters what is the problem with this sort of you know um this sort of order for me to be honest if I uh go into it my biggest problem would be uh probably it's too long maybe that'll be one of my first considerations and some other condition some other considerations would be that okay we already have the metadata sheet which already explains us what the categories are right so why do we need to put all of it here again why should we be be so repetitive so because of certain questions like this uh though there have been a lot of guides which have been created now as to how do we name our files so these are some very important tips I call them the seven tips which have been taken from the Princeton University guides and these tips are first they should be named consistently if you are following say uh language name uh place um uh participant date information if this is what you're following be consistent follow that throughout there should not be any shifting of things from one side to the other the second is that this is important so some of the archives actually want characters which are below 25 so if you actually look at this one from before this was way too long and so that is one reason why I tried to bring it below 25 um I put for the abbreviations for this because again we have the metadata sheet to fall back on right so that all of it individually can be extended or explained in the metadata sheet hence again reiterating why the metadata sheet is so important the third point is avoid special characters or spaces in a file name extremely important uh they will not be accepted use capitals and underscores instead of periods of spaces again very true use State format which I already said this is the conventionalized date format include a version number um write down the naming convention in the data management uh in the data management plan so if you are um say using the abbreviations write it down just know what your larger plan is uh the next part of it is um how do we arrange our files now that we know how to name it how do we arrange it so going back to what I had spoken about collections to sort of um recapitulate there were the uh um assemblages right where all of it was there right all your data photos videos everything um photos which are blurred photos which are not blurred etc etc from there we curated The Collection right and the collection is what goes into um the archive now the collection for the contains folders the folders for the contain media files and this is largely this is largely how the convention is and we have two structures one it's current one is the flat structure which actually a lot of in fact many digital uh archives and repositories have a flat structure uh mostly the digital repositories they have a flat structure in which folders may only contain media files and cannot contain other files whereas there are certain archives which do also accept nested structure so what is the difference between a flat structure and a nested structure and why is it important let's look at it briefly here um this is what is a flat structure it's very simple for instance you have the collection here we are going back to the Gustav Clint example we wanted to uh uh we wanted to locate a word titled degenerate art in the collection and that is our soul mode so what do we do we go into archives.org a very well known online archive and here we see that these are the collection catalogs we what we do is we click on the first one which is underlying the LACMA catalog the moment that we click this we are taken into the folder here which has all of these um all of these media files here you can see right but as I already stated I want degenerate file degenerate R sorry so what I do is within the folder I click into degenerate file degenerate art sorry I don't know why I was saying degenerate fund degenerate art and this takes me to this particular media file it immediately opens up foreign it's just three steps and we Google it uh we go into the collection we open the folder we choose what we want but say for instance going back if it was a nested collection and Within labma catalog I want to degenerate art but then I I'm presented with a lot of other things of uh Gustav klit for instance I'm giving bits all the paintings that he has done all the pictures all the video files the audio fancy there are all of it and they have been nested they are all there together that would just mean that for me to find a very particular thing that I'm looking for would be a time consuming and two I might have to open for the folders and for the folders and for the folders in order to reach something that I've been that uh that could just very easily be cataloged like this the same kind of files being put together in a sequence so yeah this is what is um a non-nested for that matter uh flat structure as we've seen here and now uh we will go back to our original idea or the original schema of what we've been doing so far we should take a break I know that it's been a lot of information a lot of things to process so let's take a break let's take a breather let's go back to where we had started we started that we wanted to achieve this we have our data we wanted to make it good enough through primary data and through secondary data so that we could make it uh good enough for an archive so that we can archive it now within the primary data we looked at a couple of things we looked at the structure of the primary data um then we also looked at how to name and store primary data so that it's it can be easily uploaded to an archive and in terms of the secondary data we looked at the basics of metadata how to fill in a metadata sheet a metadata file and now what we would look at is the process of annotation through Elan so annotation is a very important part of it right so annotation is in a non-linguistic definition or even in linguistic definition uh those are just scribbles and those are just uh um you know notes that you have so transcriptions Etc all of them come under annotation and then Elan is one of the easiest and one of the most widely used annotation tools and um unfortunately because we had to pack in a lot and it's been it's it's supposed to be an introductory um uh Workshop um we will not get into the complete depths of Elan but then I would like to show the very basics of it so that somebody who wants to go into a community or for that matter their community and they want to you know collect data and then um and then make it and then annotate it make it um make it good enough for an archive what would they do so I can show the basics but then I'll show the basics for now but then if you have any questions or whatever you can always ask me for sure after it so give me just a second I will um try and share it um okay okay is is this visible to everyone Anna is this is your blanket on screen yes so Elan always starts like this and uh and uh to be honest I was just thinking about my initial times in Milan because most times I was like why is it blank there should be something like what do we do what did we do around well this is how Elan is they give you all of the space in order to uh do what we'll do now which is we will press on new um the reason why it's not an extensive Workshop is also because um not everyone here is trained in formal Linguistics and I presume not everyone has Elan uh so I would just briefly show you um how it is done so I clicked on file I opened the new and this is a folder so there are a lot of folders you can choose it from your desktop you can choose it from your library um you know you this is an additional folder I have and so on and so forth but now I've already chosen my folder spiti So within spitty I have say all of these uh further nested files but then I will just go for this so I click on this and then what I do is I shift the file here and there is a reason why I do this why I don't directly open from here because as you see here there is an option of a template so what can be done in Elan which makes it of course very easy hence is that you could have uh say you are working on 1000 sentences that you have to transcribe you can make a template for one right and all the 1000 sentences say are about person marking and then you can make one template and then you could reuse that here you could add you could also select you could select your new file you could select your template file and then you know you don't have to make the tiers and the types uh every single time you use alarm but because for now we don't have a template we'll quickly see how that's done for now I've chosen this and I click on OK okay this yes so what we have here then is this we have um you see the recording the recorded waves here uh to give you a brief introduction into what just happened we imported the audio file here right these These are where our tears would come below um uh there are a couple of keys here if you can see this if you press on this this will play the whole file we have a similar over here one under selection wherein if you see select something let me just go to say a random part in the file just to show it very quickly um let me go to uh say I'll go to the end and okay what looks cheese and here can I see your wave somewhere here let me see is there something here so we can try listening to it this is the selection say like how I've selected this right we could try playing this yeah there is something there but then it is um not super clear let me try and find if there is something which is probably a little clearer um okay I'll just go back a little bit and see if there is something yeah so because we see that here there are something right we see a little bit of things here okay now let's check okay there is something which is a lot more clear foreign can everybody hear me uh can everybody hear what's being said here is it Audible okay so here we hear Bahar right so what we can do here is that um this is the space of our uh transcription this is where we will transcribe our data but before we get into transcribing what we have to do is that we have to create a type and a tier so what really is a type in it here and how do we really create it let's check so uh people who are say um in the Aryan speakers would probably know what is being said here it is it says Bahar which means outside which is uh uh post position in Hindi uh that we use so now that we know that it's a post position let's see what we can do is we can go into type uh we'll add a new tier type we will write the type name as uh post position okay wrong spelling post position [Music] um and well we will not get into control vocabulary etc for now uh none of this and we just add this okay so here we see that we have a post position here we have a default here as well but uh so generally what we do is we delete the default here nobody really knows why that default here exists it has no work it has no purpose whatsoever but because we are running a little behind times we will not engage in deleting the default here but generally we do delete it so we just click on ADD now oh okay so this says this because I already have uh something called post position so let me just call it probably um um I don't know position yeah just yeah let me just call it position for now just to make sure there are no uh nominees okay so now we have something here called uh position and we also have uh post position which we could uh delete it but for now let's just keep it here we'll see what happens with that so we add this oh God okay uh we might have to just delete this uh okay I will get into this again wait I will add new to your type uh so the issue is that because I had already created these uh I have to first delete it off the tier and then I'll have to delete it off the type but okay let's try this again because I'd already I'm so sorry while I was trying it I had already created uh this particular tier for sorry this particular type for Bahar it's coming in again I'm really sorry but then um for me to delete it with me and I'll have to go back to tier first delete the tier and then only I can delete the type but for now let's just keep it here there is a particular uh type called post position let's just keep it to that and then we'll see at what's at here and if we check here a tier comes under a type right so now see I write the tier name the part is the uh the tier name would be um out say I'm sure even this will say that it's already it's it already exists because I had already created it I should have just taken a different name the participate the participant is say um abha right the annotator is anujima the parent the parent here is uh none the tear type is now we have two gear types because our created post position and I then created position right but let's say we put it under uh post position what happens then see so here we get this okay now this one comes we have uh out here which is uh the tier then we have the type which is post position the participant is abha the annotator it comes here what we could also do is we could also uh arrange the tiers depending on uh color preferences right so we could also uh for instance see if we want the entire tier uh as one particular color we could also arrange it as say uh I don't want the green there but I won't say the blue there I could also do that so um and if I want to highlight the tier I can also do that and so highlighting that here depends on again if you want say we have two speakers and I want abhas tier to be highlighted in one particular color which is this and see the second is Ashok who's speaking and I want ashok's to be in a different color we could also do that so that there is a difference between um the speech of ABBA and the speech of a show but for now let's see how this goes apply um say we do it all tears with the same participant apply and let's see what happens uh okay I'm so sorry this one um I'm so sorry for this cute thing um do you think Anna there's enough time for me to uh highlight another section and then do it or do you think I should just briefly wrap this up because the Bahar part I had already done it that's why it's showing me that I should keep changing the tier and the type do you think I can choose another part of um of it from Elan should I show it or what do you suggest to you yeah if you think it's a useful demonstration go for it but I think that Elon also takes a lot of practice and this is a good example for everybody that things go wrong in Elan and it's just normal yeah I mean um yeah yeah I mean to be honest the only thing that happened here is that I should have deleted my existing type and tear and right because it is nested so a cure is nested within a type so one has to first delete all the tiers and then the type and then uh that is what I should have done but I did not do that I'm really sorry but I'm trying to so uh yeah so basically what would then happen would be okay let's just mean just let me give me a next let me just give an example so because we already created the tier out here under type right uh we just don't have a parent here here which is okay but we have a tear type as you can see which is post position which has already been created uh here this the selection which is we can just double click it and uh we can just write transcribe it as for instance right and then we click anywhere outside and this selection stays so uh this we could also do for longer files right um here on the left hand side once the default is gone here on the left hand side say if you have three speakers you can make three speaker tiers if you have say uh if you want to differentiate it based on uh if your work is on post positions for instance and you have uh in you're working on a language which is out in corner right you can also make three tiers depending on that and so on and so forth so whatever here you make is completely based on the kind of work you do so for now just to show you one in out I have Bahar here and okay this is uh not the correct um transcription I'm so sorry but yeah I just presumed it's uh the correct one and to quickly show you how then we export it what we do is file exporters we export it as a Interlinear text we click on this and what we find here is this we have the uh uh the word which is being transcribed and here it shows the time duration at what duration in the entire data has power been spoken and because we are talking about very very long audio video files right that is why this is so important so that this is just one of the entries right but suppose you have uh an entire interview that goes for an R it will be uh it will be extremely crucial to know within what time duration what particular um post position was spoken say if I'm working on post positions um there are a couple of options here uh right that you can see one can always check out I usually prefer Showtime code because um the data that that is time aligned is actually very very helpful you can also see what happens when we show uh to your left uh labels when we see here we apply changes see this is what comes the default comes because we didn't delete people and this is the tier which is out and we have uh Bahar and the time duration so there are a lot of options here where one could depending on the preferences choose apply you can also uh edit it depending on your line spacing and so on and so forth so uh this is what we do and then we can um save it in a document save it in the desktop and then later on we can also open it as a Word document wherein it will show us we could import it or open it as a Word document where in this will um this will come out as an annotated sheet so this is um the very basic offalon I'm again extremely sorry because I was practicing before um right before this and um yeah I completely forgot to delete the tears and the types but I hope irrespective that we all got a very very brief idea if you're want there can of course be another um another uh session on Milan this can that can also be a practice session which could be longer one can also talk about control vocabulary Etc they're a lot of other aspects but we can always get back to it this is supposed to be very introductory so uh just to give you a very brief idea so going back to where we were in the slide uh we spoke about the primary data secondary data the structure the basics of metadata uh The annotation tool which is learn and we did a brief alarm tutorial and uh just to sort of uh bring it all together in a step-by-step plan or a step-by-step workflow uh suppose you are say either an outsider getting into a community going into entering a community or you are somebody from the community it is very important that we have this workflow in in the head the first step is to plan uh generally you know um what are we going to um document what are the kinds of questionnaires we'll do so on and so forth the second is to write down the key metadata the third is to make a recording and make sure that we move the raw data to the computer we rename the data using the file naming convention that we just spoke about depending on uh if your goal is uh um large if you want to uh put it up in the archive um and you want to keep you make your work easier then one can follow the sequential um uh naming convention um or if you want to just store it for now you can use um you can use um um you can you can use uh just you know the numbering tradition and whatever makes sense uh you make it's very important to make sure that there is a good backup in the external hard drive uh it is highly recommended that while you're still in your field work you fill in the metadata sheet uh which you must have already gotten by the time you are in your field work um along with that uh we've already learned very little of Allah and how alarm works and how data can be processed using Elan how one can use it for transcription and translation make sure to back up the process data as well and uh and eventually archive all the data with proper file name so this is more or less very roughly you know in this um in this Workshop that we have tried to see that we have tried to go through um so trying all the ends together so there are three large uh uh functions of three large coals two um archiving uh or two uh documentation and what you do of the data is one we have to look at the accessibility of the data uh how accessible the data is uh that can be seen through metadata how uh well the metadators and the metadata gives us accessibility into our primary data uh the second thing is how discoverable the data is uh how well can it be uh navigated depending on whether it's um uh what kind of archive it is whether the um belief in uh sequential Arrangement or nested Arrangement and so on and so forth but then irrespective it's very important that the data is discovered and it's not you know once you collect it to make sure that it's not um kept uh in a hard drive or somewhere where say 100 years down the line it cannot be accessed by the community the third thing is the functionality of it and um how the data can be revived retrieved how the data can be used this also brings in a lot of things which we didn't describe or discuss now which we could again uh take it up later which uh includes uh the legality aspect of it for instance um you have um the copyright uh the issues of copyright how do we know about copyrights how do we negotiate copyright so that is another complete aspect of both metadata as well as Publications and Community work which we could always take up of course but then these are the three key aspects through which it's very important um to make sure that your data is accessible it's just it can be discovered and that it's uh that it has high functionality when it's needed um we follow Austin 2021 in considering the role of archives to be appraising materials that is collecting selectively based on a stated goal preserving those materials making known their existence and facilitating their appropriate distribution so when we speak about distribution distribution again as I pointed it out here circling back to what we started with prioritize Community access to the data it's very important to do that when we speak about facilitating their appropriate distribution distribution not just ends with the research Community but more importantly it also has to go back to the community and we have to make sure that we negotiated with the community we make sure that uh we prioritize uh the fact that they can access it in a language which is comfortable to them and uh these are some of the links uh I didn't put in a bibliography because I have largely um uh cited whatever wherever I've taken from but then these are some very important links that one can look into um I can come back to it talk about it more um if you want but then um for instance archiving for future is an excellent excellent um tool um then there is open Humanities there is um archiving and language documentation from the Cambridge handbook these are again these are again very uh these are again uh resources that are available online for free and uh finally uh that's me um circling back to where I started that's a picture of me disrupting their everydayness and yeah and then the burden definitely lies on each one of us who are um documenting a lesser known endangered language to make sure that um the community has complete access and rights over the data um yeah that's about it thank you and I'm anujima and I hope this was at least a little helpful um yeah thank you all right thanks so much onajima let's have a big round of applause for that extremely informative presentation and uh I think there are some questions in the chat in the Q a if you feel like taking those uh okay and if you want I will pull them out of the chat and paste them in so you can see them easier of um okay I'll just go on it from The Talk just give me a second uh okay there is a question uh by uh this fine nagesh by you I'm sorry if I'm not uh pronouncing your name correctly um he asked me if the version of alarm matter in archiving I mean uh to be honest I've been using my version 5.7 and even though there are new versions now I've been using this version uh I would say a piece for the past five years and I've personally not faced any issues right and um I think it this is the thing with versions right there are always some um some betterment a little bit here and there some Bots that being that are being taken care of but uh by and large I think at least for me personally uh this is the version that I've been very comfortable with and something that I've been using for years and this has worked uh perfectly for me did that answer your question yeah okay so we have Rosalind Mirasol she asked me would it be possible to share with us Elan can we download it from a secured side thank you thank you yes you can actually uh download um alarm and I think Ronnie Harris also uh gave us thank you so much for that he also immediately uh given the um link from where it can be downloaded so the MPI archive is an excellent MPI as the Max Planck Institute the archive is excellent for a lot of such resources um yeah and uh the next yeah that is so true uh the Spy right the Spy that is so true I mean I also felt like there was so much that had to be sort of put in together in an R because Elan in itself to go through it properly takes an R right if we have to understand Elan if we have to go through the steps Etc it takes a really really long time and because we did three broad things here um it was rushed I'm so sorry of course as I have been saying if you all want um more sessions later on um that can definitely be arranged and I would love to do it for sure yeah Ronnie says way better than copying tasting type going into self-made spreadsheet slowly yeah for sure and also in self-made um um spreadsheets for instance what why I love Elan is because it gives it to you down to seconds so that becomes um for instance uh difficult right sometimes to just like um figure that out during a speech or depending on however you are creating your metadata yeah but Ronnie thanks a lot as Anna said thanks a lot uh for being um uh so quick with the answering and the questions thank you so much for being so interactive in the first place I think we get a bit of an understanding of Elan it does seem very flexible okay thank you so much Robert I I still feel like it was a bummer that um I forgot to delete the type and the tier that would have made things way more beautiful but then uh we always um there's always a next time uh the Spy says to be only accept use Excel sheets for metadata how about the best software for metadata with software respects for metadata nowadays is there any change development of metadata software okay so um that's why um again to be honest metadata in itself could be a two-hour long session right because there's so much happening in the space um there is um there is for instance c-more which is doing a lot into bringing together um together a lot of platforms trying to also make metadata easy there is La meta also which is doing things so you know depending on uh the what metadata convention you're talking about uh um and what sort of metadata again right we have we already saw at least two kinds of metadata um for instance write-based metadata uh or uh copyright metadata would be completely different that would be the administrative metadata so in order to collate and bring together all of it Excel is sort of um something that has been used for years now right even though there is a transition that's been coming out um that sorry that has been coming in um but yes we should definitely talk about metadata probably have another session where we can deal with it individually look at softwares probably learn a little bit more about how samor is trying to be more integrative and so on and so forth um thank you Rosalyn thank you Katie I already okay thank you Watson um yeah Elan is free and you can download it from MPI uh there is a question I think this is an administrative question to uh Anna if he can have a separate session on alarm probably we could um is there a standard way to see through the symbols and the fonts and the archive so um this is as I already said in the beginning it's very important to decide on the archive or the uh um repository in the starting and the start right because every archive and every uh not every every is a blanket soup when most archive and most uh repositories have their own conventions and how or what is accepted according to them so um this it's generally the um so it's generally the archive or the um repository that will already tell you what are the symbols or what are the pawns or whatever what is used but then the seven uh tips that I spoke about those are largely Universal apart from the less than 25 character um specification because most archives actually don't have that but some archives too and that's why we put importance on that but largely it will be archive and repository dependent so that's why one should figure that out first and then start building the metadata sheet and then figuring out how should one go ahead with um the data my Elder was asking about some tapes that have black mold on them uh okay um is that okay this is [Music] um not really something that is my forte this question um by Tara Rose my Elder was asking about some tapes that have black mold on them is there any way to save them uh to be honest uh I wouldn't really know about it and I I mean if somebody in the audience would be able to answer this um amazing please help Tara here or Anna if you know how to save um some tips with moles would you know how to do that I I know that it theoretically is possible depending on how bad the mold is and how long it's been there um mold can accelerate the magnetization of audio tapes and if they've been sitting around long enough that they're moldy they may already be very demagnetized but there are archivists and like audio Specialists who do offer I've never had to do this they they can like vacuum the mold off very carefully and see if it's still magnetized so your local University may have Library technicians or archivists who can help you with that that's probably where I would ask yeah so we always have a person who we run to when it gets super super technical for instance like this uh oh okay oh then it's very okay so uh yeah I would suggest then um you should ask uh your University um uh technician I'm sure he'll be able to help out uh in this regard I'm not very sure if a linguist can help out with this of course uncle and unless they have a previous experience personal experience in removing mold but uh I would suggest if you have a university technician that would be a very uh they would be a very good starting point uh yeah I think we are uh the end of our questions here do do we have more I think that's it all right but if anyone has questions going forward feel free to post them in the Facebook group or anojima are you open to having people email you yeah yeah absolutely I can put my email ID on the chat box perfect and um please let me know if you have any questions I would love to mail you back um so so that is my email I did yes gmail.com all right well thank you again so much adojima this was a wonderful presentation awesome tough stuff but you made it so accessible thanks so much and uh next week we will have a slightly different format we will be meeting an hour later than usual but we will be hearing from Stephanie Witkowski at 7 000 languages about some free tools to build dictionaries including talking dictionaries and learning courses and other language materials so we'll see you next week can't wait and have a good night or morning or afternoon everybody oh and energy I think you put in your email just to the host and panelists so I will copy it back in okay okay uh okay thanks oh and I think you had like a special character G so everybody if you copy that email address it should just be a plain normal G not a fancy ipag okay I will try to rewrite that again I completely lost that it was only for uh panelists yeah two times yeah Zone sometimes can be so difficult to navigate yeah just like a lot and one thing goes here and there and you have to delete yours and tires and the choice of Technology there it is all right there it is everybody okay grab on a dreamer's email address she will be happy to answer any questions and thanks again take care everybody till next week thank you bye
Up Next

ELAN Tutorial 1: Set Up Tiers and Annotate Sign Language Data
@EI435LingASL
42.4K views•2011-06-22

Conversation Analysis: Key Concepts & Research Domains in Linguistics
@pointstoponder5186
9K views•2020-12-30

Forensic Linguistics: How Language Solves Crimes | PBS
@pbsstoried
1M views•2024-01-25

Accent Expert Explains U.S. Regional Dialects | Part 1
@WIRED
9.3M views•2021-01-21
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Linguistics









![[Introduction to Linguistics] Phonetics, International Phonetic Alphabetic, and Sound Classes](https://i.ytimg.com/vi_webp/jBpC4cftHs4/maxresdefault.webp)



























