Generalization Bias in Large Language Model Summarization of Scientific Research
2025/03/28 by Uwe Peters, Benjamin Chin‐Yee, Benjamin Chin-Yee +2 · 12 voices · 34 citations
Computer Science · Medicine · Social Sciences · #AI in Service Interactions #Artificial Intelligence in Healthcare and Education #Topic Modeling #cs.CL #cs.HC
paper · pdf · doi:10.48550/arxiv.2504.00025
openalex publication_date 2025/03/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Artificial intelligence chatbots driven by large language models (LLMs) have the potential to increase public science literacy and support scientific research, as they can quickly summarize complex scientific information in accessible terms. However, when summarizing scientific texts, LLMs may omit details that limit the scope of research conclusions, leading to generalizations of results broader than warranted by the original study. We tested 10 prominent LLMs, including ChatGPT-4o, ChatGPT-4.5, DeepSeek, LLaMA 3.3 70B, and Claude 3.7 Sonnet, comparing 4900 LLM-generated summaries to their original scientific texts. Even when explicitly prompted for accuracy, most LLMs produced broader generalizations of scientific results than those in the original texts, with DeepSeek, ChatGPT-4o, and LLaMA 3.3 70B overgeneralizing in 26 to 73% of cases. In a direct comparison of LLM-generated and human-authored science summaries, LLM summaries were nearly five times more likely to contain broad generalizations (OR = 4.85, 95% CI [3.06, 7.70]). Notably, newer models tended to perform worse in generalization accuracy than earlier ones. Our results indicate a strong bias in many widely used LLMs towards overgeneralizing scientific conclusions, posing a significant risk of large-scale misinterpretations of research findings. We highlight potential mitigation strategies, including lowering LLM temperature settings and benchmarking LLMs for generalization accuracy.
Citations
Cited by
Discussions
- LLMs suck at lit review, and they're getting worse at it. Transparently obvious from today's news (HHS's nonsense MAHA report!) and backed with recently published peer-reviewed research by Peters & Ch [bsky, 11 points, 2 comments]
- New study finds that AI chatbots overgeneralize when summarizing scientific studies: doi.org/10.1098/rsos... What made me do a spit-take is that asking them to be accurate—"do not introduce any inacc [bsky, 5 points, 0 comments]
- Good qn. I've refereed many ML models for EM...90% very poor. Generative AI may be getting worse in areas that matter (see doi.org/10.1098/rsos...). But - if they can help write good clinical notes [bsky, 2 points, 0 comments]
- Un nuevo estudio revela que los #LLM actuales rinden peor que versiones anteriores al utilizarse para resumir investigaciones Estos modelos de #IA #AI fueron casi 5 veces más propensas que los human [bsky, 2 points, 0 comments]
- Not sure why it defaults to this pic doi.org/10.1098/rsos... [bsky, 2 points, 0 comments]
- LLMs are getting worse at summarizing scientific data: “Generalization bias in large language model summarization of scientific research” doi.org/10.1098/rsos... [bsky, 1 points, 0 comments]
- TL;DR. ChatGPT, can you summarize? "Generalization bias in large language model summarization of scientific research" by Peters and Chin-Yee (2025, RSOS) doi.org/10.1098/rsos... [bsky, 1 points, 0 comments]
- "Notably, newer models tended to perform worse in generalization accuracy than earlier ones. Our results indicate a strong bias in many widely used LLMs towards overgeneralizing scientific conclusions [bsky, 1 points, 0 comments]
- **Generalization bias in large language model summarization of scientific research** “_Our results indicate a strong bias in many widely used LLMs towards overgeneralizing scientific conclusions, posi [mastodon, 0 points, 0 comments]
- We are learning more & more about large language models, like ChatGPT — including their biases. Indeed, when asked to summarize the scientific literature, many AI models are biased; they tend to extra [bsky, 0 points, 0 comments]
- Generalization bias in large language model summarization of scientific research doi.org/10.1098/rsos... [bsky, 0 points, 0 comments]
- doi.org/10.1098/rsos... < user beware [bsky, 0 points, 0 comments]
Related