vix.ing · top · new · best · stats · spec

The Corpus of Contemporary American English as the first reliable monitor corpus of English

2010/10/27 by Mark Davies, M. Davies · 3 citations
Arts and Humanities · Computer Science · Social Sciences · #Lexicography and Language Studies #Linguistic Variation and Morphology #Natural Language Processing Techniques

paper · doi:10.1093/llc/fqq018

crossref issued 2010/10/27 · crossref published 2010/10/27 · crossref published-online 2010/10/27 · openalex publication_date 2010/10/27 · crossref created 2010/10/29 · crossref published-print 2010/12/01 · crossref deposited 2019/03/05 · openalex created_date 2025/10/10 · crossref indexed 2026/08/01 · openalex updated_date 2026/08/01

Abstract

The Corpus of Contemporary American English is the first large, genre-balanced corpus of any language, which has been designed and constructed from the ground up as a ‘monitor corpus’, and which can be used to accurately track and study recent changes in the language. The 400 million words corpus is evenly divided between spoken, fiction, popular magazines, newspapers, and academic journals. Most importantly, the genre balance stays almost exactly the same from year to year, which allows it to accurately model changes in the ‘real world’. After discussing the corpus design, we provide a number of concrete examples of how the corpus can be used to look at recent changes in English, including morphology (new suffixes –friendly and –gate), syntax (including prescriptive rules, quotative like, so not ADJ, the get passive, resultatives, and verb complementation), semantics (such as changes in meaning with web, green, or gay), and lexis––including word and phrase frequency by year, and using the corpus architecture to produce lists of all words that have had large shifts in frequency between specific historical periods.

Cited by

Related