2025/11/07 by Germani, Federico, Spitale, Giovanni
Computer Science · Medicine · Social Sciences · #610 Medicine & #Artificial Intelligence in Healthcare and Education #Computational and Text Analysis Methods #Topic Modeling #health
paper · doi:10.5167/uzh-280397
openalex publication_date 2025/11/07 · openalex created_date 2025/12/11 · openalex updated_date 2026/07/28
Large language models (LLMs) are increasingly used to evaluate text, raising urgent questions about whether their judgments are consistent, unbiased, and robust to framing effects. Here, we examine inter- and intramodel agreement across four state-of-the-art LLMs tasked with evaluating 4800 narrative statements on 24 different topics of social, political, and public health relevance, for a total of 192,000 assessments. We manipulate the disclosed source of each statement to assess how attribution to either another LLM or a human author of specified nationality affects evaluation outcomes. Different LLMs display a remarkably high degree of inter- and intramodel agreement across topics, but this alignment breaks down when source framing is introduced. Attributing statements to Chinese individuals systematically lowers agreement scores across all models and, in particular, for DeepSeek Reasoner. Our findings show that LLMs’ own judgment of agreement with narrative statements exhibit systematic bias from framing effects, with substantial implications for the neutrality and fairness of LLM-mediated information systems.