AntConc 4.2.4 Tutorial: Corpus Analysis for Textual Research

Added:

Setup & Data Prep
Interface & Corpus
Keyword in Context
Plot & File View
Clusters & Collocates
Word Frequency
Keyword Analysis
Corpus Q&A
Advanced Tools

Setup & Data Prep

0:00
Playing Section
  • 1

    Guide to downloading and installing AntConc on Windows and Mac.

  • 2

    Recommend saving text data as .txt files to avoid extraction errors.

  • 3

    Process of manually compiling a corpus from a website is tedious.

Basic principles of corpus linguistics, including understanding what a corpus is and how it differs from traditional qualitative text analysis.
Fundamental linguistic terminology, specifically the distinction between 'types' (unique words) and 'tokens' (total running words), as well as 'lemmas'.
Familiarity with digital text file formats and data preparation, such as saving files as plain text (.txt) and understanding character encoding standards like UTF-8.
The conceptual meaning of 'frequency' and 'concordance' in the context of analyzing written or spoken discourse.
Utilizing Regular Expressions (Regex) within AntConc to perform complex, pattern-based linguistic queries.
Part-of-Speech (POS) tagging and lemmatization of corpora to enable more advanced syntactic and morphological searches.
Deep dive into the statistical measures of collocation and keyness, such as Mutual Information (MI) scores, T-scores, and Log-Likelihood.
Applying corpus analysis techniques to specific research domains such as Critical Discourse Analysis (CDA), stylistics, translation studies, or language pedagogy.
Transitioning to programmatic corpus linguistics tools using Python (e.g., NLTK, SpaCy) or R for large-scale, automated text mining and visualization.
3K views45likes1:15:23@keemanxpOriginal Release: 2023-11-29

This workshop demonstrates how to use AntConc 4.2.4 for corpus or textual analysis, covering the complete workflow from installing the software (with installer vs portable versions for Windows/Mac) to preparing and loading corpora, followed by six key analytical tools: KWIC (Keyword in Context) for examining word usage patterns, Plot for visualizing word dispersion across files, Cluster/N-Gram for identifying word combinations, Collocate for finding associated words, Word Frequency for counting word occurrences, and Keyword/Keyness for comparing specialized corpora against reference corpora to identify distinctive vocabulary.