vix.ing · top · new · best · stats · spec

Sensitivity of dispersion measures to distributional patterns and corpus design

2026/07/03 by Lukas Sönning, Jesse Egbert · 1 voice
Computer Science · Social Sciences · #Computational and Text Analysis Methods #Natural Language Processing Techniques #Topic Modeling

paper · doi:10.1075/ijcl.25008.son

openalex created_date 2025/10/10 · openalex publication_date 2026/07/03 · openalex updated_date 2026/07/16

Abstract

Abstract Recent work has shown that dispersion measures respond to multiple features in the data: Juilland’s D varies systematically with the number of corpus parts, and all commonly used indices are affected by the frequency of an item. This study uses a simulation approach to provide further insights into the sensitivity of dispersion measures to differences in corpus design (number of texts, average text length, distribution of text lengths) and distributional milieu (frequency and evenness of distribution). Our results suggest that, within the settings covered by our analysis, the factors frequency and evenness of distribution have roughly the same impact, though there is some variation among measures. The average text length emerges as another feature that leaves its mark on the observed scores. Finally, we note that D 2 exhibits the same weakness as D — it varies with the number of corpus parts that enter the analysis.

Citations

Cited by

Discussions

Related