Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
2025/10/27 by Liwei Jiang, Jiang, Liwei, Yuanjun Chai +17 · 15 voices · 26 citations
Computer Science · #cs.CL
paper · pdf · doi:10.48550/arxiv.2510.22954
Abstract
Language models (LMs) often struggle to generate diverse, human-like creative content, raising concerns about the long-term homogenization of human thought through repeated exposure to similar outputs. Yet scalable methods for evaluating LM output diversity remain limited, especially beyond narrow tasks such as random number or name generation, or beyond repeated sampling from a single model. We introduce Infinity-Chat, a large-scale dataset of 26K diverse, real-world, open-ended user queries that admit a wide range of plausible answers with no single ground truth. We introduce the first comprehensive taxonomy for characterizing the full spectrum of open-ended prompts posed to LMs, comprising 6 top-level categories (e.g., brainstorm & ideation) that further breaks down to 17 subcategories. Using Infinity-Chat, we present a large-scale study of mode collapse in LMs, revealing a pronounced Artificial Hivemind effect in open-ended generation of LMs, characterized by (1) intra-model repetition, where a single model consistently generates similar responses, and more so (2) inter-model homogeneity, where different models produce strikingly similar outputs. Infinity-Chat also includes 31,250 human annotations, across absolute ratings and pairwise preferences, with 25 independent human annotations per example. This enables studying collective and individual-specific human preferences in response to open-ended queries. Our findings show that LMs, reward models, and LM judges are less well calibrated to human ratings on model generations that elicit differing idiosyncratic annotator preferences, despite maintaining comparable overall quality. Overall, INFINITY-CHAT presents the first large-scale resource for systematically studying real-world open-ended queries to LMs, revealing critical insights to guide future research for mitigating long-term AI safety risks posed by the Artificial Hivemind.
Cited by
Discussions
- NEW PAPER: All LLMs respond to the same prompts the same way, basically meaning all outputs will only ever be the blandest possible rephrasing of sentences If you’re impressed by Gen AI’s “creativity” [bsky, 8 points, 2 comments]
- 📚Yesterday in our reading group, Mareike Lisker presented "Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)" by Liwei Jiang et al. (2025) Paper: arxiv.org/pdf/2510.2295 [bsky, 4 points, 0 comments]
- this paper is very important: arxiv.org/abs/2510.22954 [bsky, 3 points, 0 comments]
- they seem in general to be more homogenous than people are. arxiv.org/abs/2510.22954 [bsky, 3 points, 1 comments]
- This is dumb. Here’s a study about how all LLM’s are shit at coming up with new things and all answer very homogeneously when asked open ended questions: arxiv.org/pdf/2510.22954 [bsky, 3 points, 0 comments]
- Artificial Intelligence is killing creativity, making it a more scarce resource. Congratulations @liweijiang.bsky.social ! arxiv.org/abs/2510.22954 [bsky, 2 points, 0 comments]
- Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond) [hn, 2 points, 0 comments]
- Ondertussen alweer een wetenschappelijk artikel over AI dat aantoont dat het dodelijk is voor de creativiteit en dat alles wat de verschillende modellen produceren uiteindelijk resulteert in dezelfde [bsky, 1 points, 0 comments]
- LLMs are engines for “more of the same” - they homogenise language and ideas: culturally hazardous. This paper shows massive similarities across responses from 25 models https://arxiv.org/abs/2510.229 [bsky, 0 points, 1 comments]
- Fascinating stuff coming out of NeurIPS 2025 (annual conference, Neural Information Processing Systems - very nerdy) The highest-ranked paper? “Artificial Hivemind: The Open-Ended Homogeneity of Langu [bsky, 0 points, 0 comments]
- Doing numbers on Twitter/X this week: this paper tested 70+ LLMs on open-ended prompts and found they all produce strikingly similar outputs. Worse, the systems used to improve models actively penaliz [bsky, 0 points, 0 comments]
- Pair that with this paper, which shows most models output nearly the exact same words and statements. Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond) arxiv.org/pdf/2510 [bsky, 0 points, 1 comments]
- Not just a plagiarizing environmental destruction machine! It's also a cliché machine! [bsky, 0 points, 0 comments]
- Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond) arxiv.org/abs/2510.22954 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2510.22954 "Artificial Hivemind effect in open-ended generation of #LLM, characterized by (1) intra-model repetition, where a single model consistently generates similar responses, and m [bsky, 0 points, 0 comments]
Related