Corpus linguistics is the application of computational tools to analyze large collections of authentic language data (corpora) to reveal systematic language patterns, including word frequency, collocation, colligation, correlation, semantic prosody, and semantic preference; move analysis is a discourse-focused corpus-based approach that integrates contextual analysis (social context, communicative purpose, role, cultural values, text context, and formal text features) with linguistic analysis (lexico-grammatical features, textual patterns, and text structures) to understand how texts are organized and function within specific genres and discourse communities.
Corpus Linguistics: Corpus-Based Analysis and Move Analysis
Added:hi everyone our lesson for today is all about the corpus based approach in discourse analysis at the end of the lecture you should be able to first demonstrate understanding of the nature of corpus and corpus linguistics second analyze the characteristics of each feature of the corpus based analysis in this course and third apply the framework of move analysis in this course analysis now let us first define what is a corpus a corpus with the plural form corpora is a large collection of language usually held electronically which can be used for the purposes of linguistic analysis now let us first digest this definition of corpus first it is a large collection of language which means it is a collection of texts or of authentic texts and i have here an underlined word usually because before when there is no computer software available for corpus based analysis of language people or linguists and language researchers in particular collate all words or terminologies which may now be subjected to language research using a corpus based approach and as mentioned this collection of language can be used for linguistic analysis or to put it in a word research and because of the absence of computer software before which may be used in corpus based language analyses the earliest known corpora were compiled by hand and consistent consisted of what we now know as biblical texts through the bible in the modern era an early electronically stored corpus was the brown corpus which was developed at brown university in the early 1960s and consisting of one million words other notable more recent corpora are the bank of english which was developed by co-build at birmingham university which consists of over 500 million words the british national corpus also known as the bnc has 100 million words and the corpus of contemporary american english also known as coca has over 425 million words and still growing so there's these are just examples of the actual core corpora that we have nowadays but if you would really want to conduct your own language research using a corpus based approach in analyzing it you may consider collecting all the twits of taylor swift if you are a swifty and you would want to dedicate some of your time analyzing her language use and once you were able to collect all the tweets that she had on twitter then that now may be considered as a corpus which is now ready for analysis using different computer software depending on your choice now because corpus or the use of corpus is aligned with language research let us now take a look at the term corpus linguistics corpus linguistics is the application of computational tools to the analysis of corpora in order to reveal language patterns which systematically occur in them hence the main objective of corpus linguistics is to reveal language patterns and this is not just based on empiricism or through observation but through the use of computational tools which are readily available online with this corpus linguistics is a foundation of language research the rationale for such an analysis is that on the other hand large amounts of text can be analyzed automatically much more than would be humanly possible possible manually and that on the other hand patterns may be revealed by the computational tools which may not be obvious to the naked eye corpus techniques are capable of providing information about various features of language including lexi's multi-word phrases grammar semantics pragmatics and textual features and later on i will be presenting to you the different features of analysis if you will be using a corpus based approach in your language research and so here are the features you may have your analysis grounded on the word frequency or through co-location or it could be based on the correlation or romantic prosody or semantic preference now let us take a look at the definition nature and examples of analyzing corpora using each of this features of analysis first is the word frequency this is the most fundamental feature that can be analyzed by means of corpus techniques most corpus software such as anconc and wordsmith tools which are probably the most popular publicly available packages the former being downloadable for free can produce frequency lists which can be ordered alphabetically or by frequency depending on the need of the language researcher in the analysis of text so here now is an example of a corpus based language analysis anchored on word frequency as the feature of analysis co-build was able to identify the top 20 nouns and this are time people way man years work world thing day children life men fact house kind year place home sword and end on the other hand a biology corpus was likewise identified using word frequency and according to a research this are the top 20 nouns in biology we have cell and its plural form cells water membrane food plant root molecules plants wall energy concentration organism cytoplasm animal stem structure body parts and animals the second feature of analysis in a corpus based approach is the so-called co-location when we talk about collocation it refers to the combination of lexical words with one another say for instance the word fast is likely to co-locate according to a language research with the words like train and food that's why we have fast trained and fast food but the word fast is unlikely to be found together with the words meal or sleep on the other hand the word quick is likely to be found with shower and meal as in quick shower and quick meal but not with train or food and this two examples here of identifying the collocation of fast and quick are actually products of a language research using corpus based approach the next feature is colligation it refers to the grammatical environments in which a word occurs example is based on the finding of houi who is a language researcher according to his research the word consequence has a very low likelihood of appearing as the object of a clause specifically under correlation is the so-called textual correlation which refers to the position where the word tends to occur within a discourse according to huey once again who is a language researcher every word is primed to occur in or avoid certain positions within the discourse and this is how he defined correlation another example under correlation is the word consequence which tends to occur in sentence initial position based on textual correlation as either part of an adjunct or as part of the subject on the other hand the plural form of the word consequence which is consequences does not in favor or does tend to occur in the first sentence of paragraphs the next feature of analysis for corpus based approach is the so-called semantic prosody which refers to the meaning associations that words carry with them by virtue of their typical collocations with sets of semantically related words under semantic prosody a speaker or writer may be able to have his or her personal attitude about about a particular word whether it is of positive aura or of negative aura so here is an example of a word with a negative semantic prosody or a negative aura the word is cause cos has a negative semantic prosody because it was found out that the word cause is usually collocated with words like accident cancer concern damage death disease pain problems and trouble in contrast to the word cause is the word provide which according to a language research had a positive semantic prosody or positive aura because it is typically located with words like aid assistance care employment facilities food funds housing jobs money opportunities protection relief security services support and training and now we are about to deal with the last feature of analysis on their corpus based approach and that is semantic preference in semantic preference the concern is not with the pragmatic value which was observed in semantic prosody but with sets of words semantically related according to the systems of synonymy mironymy and autonomy which are typically associated with particular registers or genres words belonging to such sets can typically be found to co-locate in corpora such sets can be given a gloss to label the semantic preference so example semantic preferences might thus be pertaining to measurement causality or history or medicine or research articles and semantic preference is thus like semantic prosody in that it refers to the meaning relations attaching to collocating sets however it does not carry with it any sense of attitudinal meaning which was observed in semantic prosody again i would like to reiterate on this when you deal with corpus based approach in analyzing the semantic preference of a corpus you need to look at a particular set which may be including synonyms mirror names and antonyms so through that you may now be able to identify the semantic preference and what will now be the application of this learning of yours when it comes to corpus and corpus linguistics i have here one of a corpus based approach in discourse analysis which is known as move analysis move analysis is a discourse focused corpus analysis and its analytic framework is divided into two we have contextual and we have linguistic which is corpus based when it comes to the analysis for the contextual analysis the components are named social context communicative purpose role cultural values text context and formal text features for each aspect of contextual analysis under move analysis there are guide questions or contextual questions for name what you need to answer when you are in move analysis is that what is the name of the genre of which this text is a part for social context you need to answer the following questions in what social setting is this kind of text typically produced what constraints and obligations does this setting impose on writers and readers for communicative purpose you need to answer what is the communicative purpose of the text for role you need to answer what rules may be required of readers and writers in this genre for cultural values you need to answer the question what shared cultural values may be required of writers and readers in this genre for text context you need to answer the question what knowledge of other texts may be required of writers and readers in this genre and finally for formal text features what shared knowledge of formal text features or conventions is required to write effectively this genre all these aspects fall under the contextual analysis dimension of move analysis however when we talk about the corpus based features you need to deal with the following linguistic aspects such as lexico grammatical features text relations or textual patterns and text structures lexico grammatical features you need to answer the question what lexical grammatical features of the text are statistically prominent and stylistically salient for text relations or textual patterning can textual patterns be identified in the text what is the reason for such textual patterning and finally text structure in this aspect you need to answer how is the text organized as a series of units of meaning what is the reason for this organization having been able to present to you these guide questions you may take note of the fact that for move analysis as a corpus based approach in discourse analysis you need to deal with not just the corpus based features of a particular text but also the contextual questions involved in such text and this is my main reference i based my discussion on corpus corpus linguistics and also the contextual questions for move analyses from the discourse in english language education book of flower jew that's all for today and thank you
Up Next

Corpus Search Tutorial: English-Corpora.org Queries
@ekbphd3200
683 views•2024-09-09

American English OKAY Over Time: A Diachronic Interactional Linguistic Study
@Abralin
1.2K views•2020-07-30

Forensic Linguistics: How Language Solves Crimes | PBS
@pbsstoried
1M views•2024-01-25

Accent Expert Explains U.S. Regional Dialects | Part 1
@WIRED
9.3M views•2021-01-21
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Linguistics












































