Corpus Linguistics: AntConc, COCA & Collocation
Learning Goal: Analyze patterns of lexical collocation, semantic prosody, and grammatical variation in natural language by querying large-scale digital corpora using tools like AntConc and the Corpus of Contemporary American English (COCA).
- Prerequisites: Basic understanding of grammatical parts of speech (nouns, verbs, adjectives, prepositions). No prior programming or statistical background required.
- Estimated Total Study Time: 12 Hours
Module 1: Introduction to Corpus Linguistics & Digital Corpora
This module establishes the foundational principles of corpus linguistics. You will learn how modern linguistics uses empirical, computer-assisted methodologies to analyze large-scale, machine-readable collections of natural language (corpora). We will demystify core terminology—such as tokens, types, and lemmas—and explore the principles of representative corpus design.
Note on Video Coverage: The discovered video pool has limited coverage of the mathematical and theoretical concepts of Token-to-Type Ratio (TTR) and lemma generation. To supplement these gaps, we highly recommend independently searching for: "types tokens lemmas linguistics tutorial" and "representative corpus design corpus linguistics".
Recommended Videos
Why this video is valuable: This video serves as an excellent, high-level introductory gateway to the field. It covers the core analytical features used in corpus-based language research, including frequency analysis and word-level distributions. It provides a solid mental framework for how quantitative data translates into qualitative insights.
Why this video is valuable: This practical guide introduces the conceptual differences between word "types" and "tokens" using software interfaces. Understanding this distinction is vital for analyzing lexical diversity (e.g., through Type-Token Ratios) in any custom-built corpus.
Module 1 Knowledge Check
- Define the difference between a "token" (total running words) and a "type" (unique vocabulary items).
- Explain how a "lemma" clusters morphological variations (e.g., sing, sang, sung, sings) under a single dictionary headword.
- Describe the baseline criteria for representative corpus design, including balance, size, and genre distribution.
Module 2: Navigating the Corpus of Contemporary American English (COCA)
The Corpus of Contemporary American English (COCA) is a premier, multi-genre digital monitor corpus containing over 1 billion words of American English. This module covers step-by-step navigation of its online interface. You will master basic list queries, search syntax for part-of-speech (POS) tagging, wildcard operators, and multi-register comparisons.
Recommended Videos
Why this video is valuable: Presented by an experienced academic, this comprehensive walk-through explains the structure of COCA, its underlying 1-billion-word data balance, and how to execute baseline searches. It is the perfect starting point to understand the overall layout of the database interface.
Why this video is valuable: This video focuses on part-of-speech (POS) tags in COCA. You will learn the exact syntax to isolate lexical functions—for example, how to target the verb break (break.[v*]) versus the noun break (break.[n*]). This is a critical skill for removing syntactic ambiguity in data extraction.
Why this video is valuable: A deep dive into searching the english-corpora.org platform. It demonstrates register-based comparisons, showing how to filter search results by genre (e.g., spoken language, academic papers, news, fiction) to track contextual variation.
Why this video is valuable: A brief, highly practical guide on advanced searches. It demonstrates how to combine lexical queries with structural grammar rules, such as searching for all verbs followed immediately by a preposition (e.g., using [v*] [prep]).
Module 2 Knowledge Check
- Execute a part-of-speech tagged query in COCA to search exclusively for object as a verb vs. object as a noun.
- Construct a query utilizing wildcards to locate all adjectives ending in "-less" (e.g.,
*less.[j*]). - Compare the frequency of a target word across "Academic" and "Spoken" registers using the Chart feature.
Module 3: Concordancing and Lexical Collocations using AntConc
AntConc is a free, cross-platform corpus toolkit for analyzing local, custom-built raw text corpora. In this module, you will learn to build a clean local database and navigate AntConc's toolkit. We pay special attention to data preparation: converting raw files (PDFs/Word documents) into clean, plain UTF-8 text files to prevent software crashes. You will then master Concordance (KWIC) searches and generate collocations.
Recommended Videos
Why this video is valuable: Created directly by Lawrence Anthony, the developer of AntConc, this tutorial explains the standard workflow of downloading the software, setting up database schemas, and importing raw texts in modern AntConc 4 releases.
Why this video is valuable: This video addresses text file cleaning, a vital step in data preparation. It demonstrates how to remove code snippets, formatting issues, and numeric junk using find-and-replace tools before running corpus software. This prevents AntConc from miscalculating word frequencies or failing to read custom corpora.
Why this video is valuable: A comprehensive, workshop-length guide to advanced feature sets in AntConc 4. It covers installation variations, indexing, and how to utilize tools like Word List, Collocates, and N-Grams to analyze textual patterns.
Why this video is valuable: A quick demonstration of the Key Word in Context (KWIC) search interface. It shows how to pull up centered concordance lines from a loaded corpus database.
Module 3 Knowledge Check
- Prepare a raw document for AntConc by converting it to plain text (
.txt) and saving it with UTF-8 encoding. - Run a KWIC search in AntConc and adjust the search window limits (left and right context span).
- Sort concordance lines alphabetically by the first word to the right (R1) of the search keyword to identify immediate grammatical collocates.
Module 4: Analyzing Semantic Prosody and Preference
Words often carry positive, negative, or neutral associations that are not visible in their isolated dictionary definitions, but reveal themselves through collocational patterns. This communicative aura is known as semantic prosody (or evaluative prosody). This module explores the distinction between semantic preference (the grammatical or semantic class of words that cluster together, e.g., words related to liquid near pour) and semantic prosody (the evaluation expressed, e.g., negative outcomes near cause).
Recommended Videos
Why this video is valuable: This video introduces semantic prosody by looking at lexical "neighborhoods." It uses real examples like rife, demonstrating how it habitually appears next to negative words like despair or disease, staining its overall evaluative meaning.
Why this video is valuable: This video clearly distinguishes between semantic preference (the semantic field of surrounding collocating words) and semantic prosody (the surrounding evaluative tone/context). It does this by analyzing lexical bundles across language databases.
Why this video is valuable: A quick, practical session analyzing the word donate. It explores how its collocations (e.g., money, blood, organs) reveal its positive semantic prosody, demonstrating how to "spot the pattern" in natural language data.
Why this video is valuable: This video demonstrates how semantic prosody works in applied analysis. It explains how political speakers use collocational "colorings" to subtly project negative or positive evaluations onto seemingly neutral topics.
Module 4 Knowledge Check
- Contrast "semantic preference" (association with a class of semantic objects) and "semantic prosody" (overall evaluative feeling).
- Use COCA's Collocate tool to determine whether the verb cause displays positive, negative, or neutral semantic prosody based on its top nominal collocates.
- Explain how a word can have a neutral literal meaning but a highly charged, biased "evaluative profile" in natural speech.
Module 5: Investigating Grammatical Variation and Change
Languages change over time (diachronic variation) and differ across genres (synchronic register variation). In this final module, you will learn how to use digital corpora to track language changes over decades and explore grammatical variations between registers, such as informal spoken language and formal academic writing.
Recommended Videos
Why this video is valuable: In this expert lecture, Prof. Anke Lüdeling explains how diachronic corpora make the study of language change systematic, transparent, and reproducible. She walks through tracking morphological, lexical, and syntactic changes across time.
Why this video is valuable: Mark Davies, the developer of COCA, explains how large-scale, multi-decade digital corpora allow researchers to track low-frequency language shifts across genres, dialects, and time.
Why this video is valuable: This video focuses on historical corpus analysis over decades. It highlights how huge text databases (like those tracking 400 million words over 200 years) reveal patterns of structural change that smaller, shorter studies miss.
Why this video is valuable: Using conversational samples from the 1960s to the present, this case study shows how researchers track the diachronic development of common expressions (like okay) in informal speech.
Module 5 Knowledge Check
- Define the difference between a diachronic corpus study (tracking change over time) and a synchronic study (comparing different registers at the same point in time).
- Use COCA's History/Chart function to track how a grammatical construction (e.g., the rise of the get-passive, like get fired vs. be fired) has changed in frequency over the last few decades.
- Explain how grammatical productivity shifts can be tracked and measured using large corpora.
Course Map
Key People Index
- Dr. Laurence Anthony: Professor at Waseda University and developer of AntConc, one of the most widely used free concordancing software tools in the world.
- Dr. Mark Davies: Former Professor of Linguistics at Brigham Young University and the creator of COCA (Corpus of Contemporary American English) and other massive English corpora at english-corpora.org.
- Dr. Anke Lüdeling: Professor of Corpus Linguistics at Humboldt-Universität zu Berlin, specializing in historical language change and diachronic corpus methodology.
- Dr. Elizabeth Couper-Kuhlen: Leading interactional sociolinguist known for her work on spoken English, prosody, and language changes over time.
Final Self-Assessment
Complete this comprehensive checklist to measure your mastery of corpus linguistics tools and concepts.
- Explain the linguistic difference between word types, tokens, and lemmas, and compute a basic Type-Token Ratio (TTR).
- Clean a raw text document by removing unwanted formatting, checking for valid UTF-8 encoding, and converting it to a
.txtfile ready for import into AntConc. - Import a custom set of text files into AntConc, configure the Corpus Manager, and run a standard KWIC search for a target keyword.
- Sort search results in AntConc by their left and right context (e.g., L1, R1) to find immediate grammar and word partners.
- Locate and interpret the statistical strength of collocations in COCA using mutual information (MI) scores.
- Define semantic prosody and explain how a word can carry a positive or negative evaluative association.
- Contrast "semantic preference" (clustering with a specific semantic category) with "semantic prosody" (carrying a positive or negative evaluative aura) using examples.
- Query COCA using Part-of-Speech (POS) tags to distinguish between different grammatical functions of the same word.
- Conduct a diachronic study in COCA to track changes in word frequency over several decades.
- Compare the grammatical patterns of a word or phrase across different genres (such as informal speech vs. formal academic prose) using the Register Chart feature.

















