vix.ing · top · new · best · stats

Do LLMs agree on multi-category journal classification? Interpreting systematic divergence in Scopus-based journal categorization

2026/07/29 by Eungi Kim, Sein Min, Jason Lim Chiu
Decision Sciences · Social Sciences · #Academic Publishing and Open Access #Computational and Text Analysis Methods #scientometrics and bibliometrics research

paper · doi:10.1177/09610006261468375

openalex publication_date 2026/07/29 · openalex created_date 2026/07/30 · openalex updated_date 2026/07/31

Abstract

This paper aims to examine how four large language models (LLMs), Claude 3 Haiku, GPT 4o Mini, Gemini Flash 2.0, and DeepSeek Chat, classify 273 Library and Information Science (LIS) journals into 85 Scopus subject categories based on journal scope statements and to determine what their disagreements reveal about disciplinary boundary setting. Perfect agreement across all four models occurred in only 2.6% of cases and was confined to journals with unambiguous disciplinary positioning. Although the models matched the Scopus-assigned LIS category at rates ranging from 72.9% to 81.7%, they differed substantially in their assignment of additional categories, with some models consistently producing broader category sets and others adopting more restrictive classification strategies. A linear mixed-effects model confirmed that these differences were systematic rather than random. The findings have two important implications. First, set-overlap metrics such as Jaccard similarity inherently favor parsimonious classifiers irrespective of substantive accuracy. Second, the marked decline in agreement from the Scopus-assigned LIS category to additional categories suggests that classification systems may benefit from distinguishing between high-confidence and low-confidence assignments. Overall, inter-model disagreement appears to reflect the inherent multiplicity of disciplinary boundary setting rather than correctable technical error.

Citations

Related