2020/01/07 by Andrés Chacoma, Chacoma, Andrés, Damián H. Zanette +1 · 4 citations
Computer Science · Physics and Astronomy · #Advanced Text Analysis Techniques #Authorship Attribution and Profiling #Computation and Language (cs.CL) #Data Analysis #FOS: Computer and information sciences #FOS: Physical sciences #Natural Language Processing Techniques #Statistics and Probability (physics.data-an) #cs.CL #physics.data-an
paper · pdf · doi:10.48550/arxiv.2001.02178
arxiv created 2020/01/07 · openalex publication_date 2020/01/07 · arxiv updated 2020/01/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We study the relationship between vocabulary size and text length in a corpus of 75 literary works in English, authored by six writers, distinguishing between the contributions of three grammatical classes (or ``tags,'' namely, \it nouns, \it verbs, and \it others), and analyze the progressive appearance of new words of each tag along each individual text. While the power-law relation prescribed by Heaps' law is satisfactorily fulfilled by total vocabulary sizes and text lengths, the appearance of new words in each text is on the whole well described by the average of random shufflings of the text, which does not obey a power law. Deviations from this average, however, are statistically significant and show a systematic trend across the corpus. Specifically, they reveal that the appearance of new words along each text is predominantly retarded with respect to the average of random shufflings. Moreover, different tags are shown to add systematically distinct contributions to this tendency, with \it verbs and \it others being respectively more and less retarded than the mean trend, and \it nouns following instead this overall mean. These statistical systematicities are likely to point to the existence of linguistically relevant information stored in the different variants of Heaps' law, a feature that is still in need of extensive assessment.