COCA Tutorial: Introduction to Corpus of Contemporary American English

Added:

Explore COCA
Wildcard Search
New Words & Cohorts
Searching for Forms
Chart Distributions
Interrogating Patterns
Lemitize & Sections
Direct vs. Indirect
Refine & Repeat

Explore COCA

0:01
Playing Section
  • 1

    Intro to the Corpus of Contemporary American English (COCA).

  • 2

    Over 425 million words with powerful search tools.

  • 3

    Navigate the interface and set up a free account.

Basic understanding of corpus linguistics, including what a corpus is and how it differs from traditional dictionaries or grammar guides.
Familiarity with fundamental grammatical categories and parts of speech (e.g., nouns, verbs, adjectives, prepositions).
An understanding of basic morphology, specifically the distinction between a word form (e.g., 'running', 'ran') and its base lemma ('run').
General familiarity with search query logic, such as using basic filters or search operators in digital databases.
Advanced search techniques in COCA, including collocate analysis, mutual information (MI) scores, and syntax-based queries.
Applying corpus data to lexicography, language pedagogy, and the creation of data-driven learning materials for ESL/EFL students.
Diachronic and dialectal comparison by examining linguistic variations between COCA and other corpora, such as COHA (historical) or BNC (British English).
Using quantitative statistical methods (such as chi-square or log-likelihood tests) to analyze the significance of word frequency distributions across different genres.
91.2K views0likes19:35@TheGrammarLabOriginal Release: 2012-07-12

The Corpus of Contemporary American English (COCA) is a free, web-based monitor corpus containing approximately 425 million words of American English text, developed by Mark Davies and his team at Brigham Young University. This video introduces the COCA interface, demonstrating how to use wildcards (asterisks with and without spaces), lemmatization (square brackets), and display options (list vs. chart) for effective corpus searches. Key concepts include understanding that raw frequency alone is insufficient for meaningful analysis, normalizing frequencies per million words for valid cross-text-type comparisons, and recognizing that graphical displays require careful interpretation. The video emphasizes that successful corpus research requires clear research questions, strategic use of search tools, and awareness of data limitations.