Corpus Linguistics for Variation Analysis | Mark Davies

Added:

Corpus Linguistics Basics
Core Methodologies
Genre-Based Variation
Dialectal Variation
Historical Language Change
Societal Discourse Analysis
Q&A on Methodology

Corpus Linguistics Basics

4:02
Playing Section
  • 1

    Emphasizes empirical data and corpus size for robust analysis.

  • 2

    Uses large corpora like COCA to study low-frequency constructions.

  • 3

    Stresses representativeness across genres for valid linguistic conclusions.

Fundamental concepts of sociolinguistics, specifically how language varies across time (diachronic), space (dialectal), and social context (registers/genres).
Basic terminology of corpus linguistics, including definitions of 'corpus', 'concordance', 'collocation', and 'frequency distribution'.
Understanding quantitative measurements in linguistics, particularly normalized frequency (e.g., occurrences per million words) to compare text collections of different sizes.
Familiarity with the purpose and structure of major reference corpora, such as the Corpus of Contemporary American English (COCA) or the British National Corpus (BNC).
Conducting advanced corpus-based research using multi-dimensional analysis (MDA) to map linguistic features across different registers.
Learning Python or R programming for automated corpus processing, text mining, and advanced statistical modeling of language variation.
Applying corpus linguistics methodologies to practical domains like lexicography (dictionary building), forensic linguistics, and critical discourse analysis.
Designing, compiling, and cleaning custom niche corpora to analyze specific dialects, internet slang, or specialized professional genres.
2.3K views98likes1:22:08@AbralinOriginal Release: 2020-05-26

Large online corpora revolutionize linguistic research by providing empirical evidence for studying historical, dialectal, and genre-based variation in language. Unlike theoretical approaches that may dismiss data contradicting hypotheses, corpus linguistics values real-world language data, enabling researchers to analyze phenomena like collocations, word frequencies, syntactic patterns, and semantic shifts across different time periods, dialects, and genres. The methodology requires both sufficient corpus size (measured in billions of words) and careful attention to representativeness across different text types to draw reliable generalizations about language use and change.