Creating a Learner Corpus with POS and Vocabulary Profiling

Added:

Corpus Setup
Manual Correction
Text Inspection
Word Profiling
List Cleaning
Level Tagging
Output Review
Final Processing

Corpus Setup

0:00
Playing Section
  • 1

    Start with plain text files for building a personal corpus.

  • 2

    Use a free tagger like Clause 7 for detailed word analysis.

  • 3

    Keep both tagged and untagged versions for readability.

Basic concepts of corpus linguistics, including what a corpus is and how it is used in linguistic research.
Fundamental understanding of grammatical categories and Parts of Speech (POS) tagging concepts.
Familiarity with vocabulary acquisition metrics, such as word frequency, lemmas, and word families.
Introduction to second language acquisition (SLA) terminology and the significance of studying learner language.
Advanced corpus analysis using concordance software (like AntConc) to study collocations, n-grams, and lexical bundles.
Error annotation and analysis methodologies to systematically categorize and evaluate learner errors within the corpus.
Applying statistical methods (using R or Python) to measure lexical diversity, density, and significance of linguistic patterns.
Pedagogical application of learner corpora, such as designing data-driven learning (DDL) materials and tailored vocabulary curricula.
307 views6likes14:40@johnboziOriginal Release: 2019-09-05

This video demonstrates how to create a learner corpus by tagging student writing with part-of-speech information using Clause 7 free tagger, then analyzing vocabulary levels using AntWordProfiler with custom CEFR-aligned vocabulary lists (A1-C2), allowing teachers to systematically profile student language proficiency and identify areas needing intervention.