2026/07/03 by Lukas Sönning, Jesse Egbert · 1 voice
Computer Science · Social Sciences · #Computational and Text Analysis Methods #Natural Language Processing Techniques #Topic Modeling
paper · doi:10.1075/ijcl.25008.son
openalex created_date 2025/10/10 · openalex publication_date 2026/07/03 · openalex updated_date 2026/07/16
Abstract Recent work has shown that dispersion measures respond to multiple features in the data: Juilland’s D varies systematically with the number of corpus parts, and all commonly used indices are affected by the frequency of an item. This study uses a simulation approach to provide further insights into the sensitivity of dispersion measures to differences in corpus design (number of texts, average text length, distribution of text lengths) and distributional milieu (frequency and evenness of distribution). Our results suggest that, within the settings covered by our analysis, the factors frequency and evenness of distribution have roughly the same impact, though there is some variation among measures. The average text length emerges as another feature that leaves its mark on the observed scores. Finally, we note that D 2 exhibits the same weakness as D — it varies with the number of corpus parts that enter the analysis.