Using Large Corpora to Analyze Language Variation and Change

Added:

Corpus Size & Scope
Genre Variation in COCA
National Dialect Differences
Cultural Insights via Corpora
Historical Change with COHA
Semantic Shift over Time
Recent Trends in COCA
Real-Time Language Tracking
Corpus Analysis Q&A

Corpus Size & Scope

0:09
Playing Section
  • 1

    Explains the role of large corpora in linguistic analysis.

  • 2

    Introduces COCA, GloWbE, and COHA as key research tools.

  • 3

    Highlights the advantage of size for studying rare constructions.

Basic understanding of corpus linguistics, including what a corpus is and how representative language databases are compiled.
Foundational concepts of sociolinguistics, specifically how language varies across different social groups, regions, genres, and registers.
The distinction between diachronic (historical/over time) and synchronic (at a specific point in time) language analysis.
Familiarity with basic quantitative concepts in linguistics, such as word frequency, concordance, and collocation.
Hands-on proficiency with corpus query tools and software, such as AntConc, Sketch Engine, or the BYU corpora interface.
Advanced statistical methods for corpus linguistics, including multidimensional analysis (MDA) and mixed-effects regression modeling for linguistic variables.
Methodologies for designing, building, cleaning, and annotating (such as part-of-speech tagging and parsing) your own specialized corpus.
Application of corpus-based variationist findings to Natural Language Processing (NLP), computational social science, and the training of Large Language Models.
628 views0likes47:42@laelwebinars9025Original Release: 2020-06-20

Large corpora enable comprehensive study of language variation across genres, dialects, and historical periods by providing sufficient data to analyze low-frequency phenomena; for example, COCA (1 billion words) reveals genre-based patterns like the increase of 'awesome' in informal contexts, GloVe (2 billion words from 20 countries) exposes dialectal differences such as 'snuck' being more common in American English, and COHA (400 million words from 1810-2019) tracks historical changes like the semantic shift of 'gay' from meaning 'bright/happy' to 'homosexual'.