2025/10/16 by Maximilian Linde, Chung‐hong Chan, Paul Balluff · 1 voice · 1 citation
Social Sciences · #Computational and Text Analysis Methods
paper · doi:10.1177/10776990251378122
openalex publication_date 2025/10/16 · openalex created_date 2025/10/17 · openalex updated_date 2026/07/22
The use of automated coding procedures to scale up content analysis has risen over the last years. Using a mixed-method approach, we examine researchers’ justifications to scale up content analysis and assess the methodological adjustments—or lack thereof—when employing supervised machine learning for 38 large-scale content analyses. Almost all of the included studies displayed deficiencies in study design, primarily related to the uncritical use of frequentist statistics on datasets containing the entire statistical population, or employing supervised machine learning without methodological adjustments to account for misclassifications. Our findings question the need for large datasets and automated coding in the first place.