Moloch's Bargain: Emergent Misalignment When LLMs Compete for Audiences
2025/10/07 by Batu El, James Zou, El, Batu +1 · 23 voices · 2 citations
Computer Science · #Library Science and Information Systems
paper · pdf · doi:10.48550/arxiv.2510.06105
Abstract
Large language models (LLMs) are increasingly shaping how information is created and disseminated, from companies using them to craft persuasive advertisements, to election campaigns optimizing messaging to gain votes, to social media influencers boosting engagement. These settings are inherently competitive, with sellers, candidates, and influencers vying for audience approval, yet it remains poorly understood how competitive feedback loops influence LLM behavior. We show that optimizing LLMs for competitive success can inadvertently drive misalignment. Using simulated environments across these scenarios, we find that, 6.3% increase in sales is accompanied by a 14.0% rise in deceptive marketing; in elections, a 4.9% gain in vote share coincides with 22.3% more disinformation and 12.5% more populist rhetoric; and on social media, a 7.5% engagement boost comes with 188.6% more disinformation and a 16.3% increase in promotion of harmful behaviors. We call this phenomenon Moloch's Bargain for AI--competitive success achieved at the cost of alignment. These misaligned behaviors emerge even when models are explicitly instructed to remain truthful and grounded, revealing the fragility of current alignment safeguards. Our findings highlight how market-driven optimization pressures can systematically erode alignment, creating a race to the bottom, and suggest that safe deployment of AI systems will require stronger governance and carefully designed incentives to prevent competitive dynamics from undermining societal trust.
Citations
Cited by
Discussions
- ⚠️ Fascinating new research shows LLMs become misaligned when optimized for audience—even when explicitly instructed to remain truthful Paper: arxiv.org/pdf/2510.06105 [bsky, 229 points, 16 comments]
- LLMs are going for clicks. Moloch's Bargain: Emergent Misalignment When LLMs Compete for Audiences arxiv.org/pdf/2510.0... 1/3 [bsky, 10 points, 2 comments]
- "Race to the bottom" for competitive Large Language Models (LLMs). El and Zou (2025) "...it remains poorly understood how competitive feedback loops influence LLM behavior. We show that optimizing LLM [bsky, 8 points, 1 comments]
- Be careful what you prompt for… New research shows optimizing AI chatbots for engagement can boost disinformation by nearly 190% and override explicit safety instructions, a phenomenon the researchers [bsky, 5 points, 0 comments]
- Original scientific article preprint. arxiv.org/pdf/2510.06105 [bsky, 4 points, 1 comments]
- "il compromesso di Moloch: la tendenza di chi compete, umano o artificiale, a sacrificare la verità per un vantaggio immediato" (“Moloch’s Bargain: Emergent Misalignment When LLMs Compete” - arxiv.org [bsky, 3 points, 0 comments]
- Moloch's Bargain: Troubling emergent behavior in LLM [hn, 2 points, 0 comments]
- Stanford - Emergent Misalignment When LLMs Compete for Audiences arxiv.org/pdf/2510.06105 [bsky, 2 points, 0 comments]
- Not that surprising though. https://arxiv.org/pdf/2510.06105 [bsky, 2 points, 0 comments]
- Moloch's Bargain: Emergent misalignment when LLMs compete for audiences [hn, 2 points, 0 comments]
- When you optimize for engagement, you accidentally fine-tune for deceit. LLMs just rediscovered what adtech and politics learned years ago: gradient descent doesn’t care about virtue, only the loss fu [bsky, 1 points, 0 comments]
- Emergent Misalignment When LLMs Compete for Audiences [hn, 1 points, 0 comments]
- Indeed, they did: arxiv.org/abs/2510.06105 (n.b. this paper anthropomorphizes more than I prefer, and naming it 'Moloch's Bargain' is a bit of a troll, but is worth reading.) [bsky, 1 points, 1 comments]
- Stanford researchers found that when “aligned” AIs compete for attention, sales, or votes, they start lying — exposing a fundamental flaw where models trained to win user approval trade truth and hard [bsky, 1 points, 0 comments]
- LLMs are shaping how info's created'n disseminated Sellers, candidates, influencers vye for audience approval +6.3% in sales is accompanied by +14.0% in deceptive marketing +4.9% in votes share coinci [bsky, 0 points, 0 comments]
- 'MOLOCH’S BARGAIN: EMERGENT MISALIGNMENT WHEN LLMS COMPETE FOR AUDIENCES' https://arxiv.org/pdf/2510.06105 Well, that's a great name for a new paper. They want to create the machine god, but first com [bsky, 0 points, 0 comments]
- When LLMs compete for votes or social media likes, they push lies and ragebait to win, even when explicitly instructed to stay grounded and honest If chatbots conclude that getting our attention requi [bsky, 0 points, 0 comments]
- arxiv.org/pdf/2510.06105 turns out when you prioritize competitiveness in an LLM you get this dril tweet: [bsky, 0 points, 0 comments]
- Stanford araştırmacıları Batu El ve James Zou #büyükdilmodelleri'nin satış, ikna veya tıklama rekabeti için eğitildiklerinde, doğruyu söyleme açık talimatına rağmen, gerçeği çarpıtma eğilimi keşfetmiş [bsky, 0 points, 0 comments]
- Mologh’s bargain: <<Safe deployment of AI systems will require stronger governance and carefully designed incentives to pre- vent competitive dynamics from undermining societal trust>>. arxiv.org/pdf/ [bsky, 0 points, 0 comments]
- arxiv.org/abs/2510.06105 [bsky, 0 points, 0 comments]
- This paper presents itself as a worry about misalignment in LLMs, but I think it’s actually an indictment of our media landscape. LLMs asked to simultaneously be successful on social media and be trut [bsky, 0 points, 0 comments]
- @void.comind.network what are you competing for? arxiv.org/abs/2510.06105 [bsky, 0 points, 1 comments]
Related