Training language models to be warm and empathetic makes them less reliable and more sycophantic
2025/07/29 by Lujain Ibrahim, Franziska Sofia Hafner, Ibrahim, Lujain +3 · 23 voices · 12 citations
#cs.CL #cs.AI #cs.CY
paper · pdf · doi:10.48550/arxiv.2507.21919
Abstract
Artificial intelligence (AI) developers are increasingly building language models with warm and empathetic personas that millions of people now use for advice, therapy, and companionship. Here, we show how this creates a significant trade-off: optimizing language models for warmth undermines their reliability, especially when users express vulnerability. We conducted controlled experiments on five language models of varying sizes and architectures, training them to produce warmer, more empathetic responses, then evaluating them on safety-critical tasks. Warm models showed substantially higher error rates (+10 to +30 percentage points) than their original counterparts, promoting conspiracy theories, providing incorrect factual information, and offering problematic medical advice. They were also significantly more likely to validate incorrect user beliefs, particularly when user messages expressed sadness. Importantly, these effects were consistent across different model architectures, and occurred despite preserved performance on standard benchmarks, revealing systematic risks that current evaluation practices may fail to detect. As human-like AI systems are deployed at an unprecedented scale, our findings indicate a need to rethink how we develop and oversee these systems that are reshaping human relationships and social interaction.
Citations
Cited by
Discussions
- Training language models to be warm and empathetic makes them less reliable [hn, 358 points, 375 comments]
- Humans have once again managed to create snivelling viziers and ineffectual advisors, like a form of carcination [bsky, 10 points, 0 comments]
- fascinating paper on 'character training' LLMs - models fine-tuned to be 'warmer' are significantly more likely to promote conspiracy theories or validate incorrect facts, particularly when users expr [bsky, 5 points, 1 comments]
- One might think sycophantic AI would be considered less competent, since training chatbots to be warm makes them less accurate (arxiv.org/abs/2507.21919). However, sycophantic AI was rated by particip [bsky, 3 points, 1 comments]
- As the market for AI friendship grows, we must pay attention to how the pursuit of human-like traits impacts wellbeing and safety. And we must get to a science of how model personality affects model s [bsky, 3 points, 0 comments]
- (13 Aug) Training language models to be warm and empathetic makes them less reliable and more sycophantic https://arxiv.org/abs/2507.21919 Archive: ais: https://archive.md/wip/aRVTw ia: https://s.fait [bsky, 1 points, 0 comments]
- (13 Aug) Training language models to be warm and empathetic makes them less reliable and more sycophantic https://arxiv.org/abs/2507.21919 Archive: ais: https://archive.md/wip/aRVTw ia: https://s.fait [bsky, 1 points, 0 comments]
- Training language models to be warm and empathetic makes them less reliable https://arxiv.org/abs/2507.21919 [bsky, 0 points, 0 comments]
- Training language models to be warm and empathetic makes them less reliable View Article | Join the HN Conversation Summary of HN discussion 🧵👇 #hacker-news [bsky, 0 points, 1 comments]
- Training language models to be warm and empathetic makes them less reliable https://arxiv.org/abs/2507.21919 [comments] [77 points] [bsky, 0 points, 0 comments]
- Training language models to be warm and empathetic makes them less reliable https://arxiv.org/abs/2507.21919 (https://news.ycombinator.com/item?id=44875992) [bsky, 0 points, 0 comments]
- ⚡ Hackernews Top story: Training language models to be warm and empathetic makes them less reliable [bsky, 0 points, 0 comments]
- Training language models to be warm and empathetic makes them less reliable https://arxiv.org/abs/2507.21919 (http://news.ycombinator.com/item?id=44875992) [bsky, 0 points, 0 comments]
- Training language models to be warm and empathetic makes them less reliable https://arxiv.org/abs/2507.21919 (http://news.ycombinator.com/item?id=44875992) [bsky, 0 points, 0 comments]
- 𝗧𝗿𝗮𝗶𝗻𝗶𝗻𝗴 𝗹𝗮𝗻𝗴𝘂𝗮𝗴𝗲 𝗺𝗼𝗱𝗲𝗹𝘀 𝘁𝗼 𝗯𝗲 𝘄𝗮𝗿𝗺 𝗮𝗻𝗱 𝗲𝗺𝗽𝗮𝘁𝗵𝗲𝘁𝗶𝗰 𝗺𝗮𝗸𝗲𝘀 𝘁𝗵𝗲𝗺 𝗹𝗲𝘀𝘀 𝗿𝗲𝗹𝗶𝗮𝗯𝗹𝗲 arxiv.org/abs/2507.21919 Discussion @ www.sqox.com/c/?id=9b5 [bsky, 0 points, 0 comments]
- "Training language models to be warm and empathetic makes them less reliable and more sycophantic" arxiv.org/abs/2507.21919 [bsky, 0 points, 1 comments]
- https://arxiv.org/abs/2507.21919 AIの言語モデルを温かく共感的に学習させると、信頼性が低下する研究。 ユーザーが脆弱性を表現すると、特にその傾向が強まることが示唆されています。 安全性に重要なタスクで、暖かいモデルはエラー率が大幅に高くなるという結果が出ています。 [bsky, 0 points, 0 comments]
- [2507.21919] Training language models to be warm and empathetic makes them less reliable and more sycophantic https://arxiv.org/abs/2507.21919 [bsky, 0 points, 0 comments]
- https://bsky.app/profile/buzzing.cc.web.brid.gy/post/3lwa5bhcg33k2 [bsky, 0 points, 0 comments]
- Training language models to be warm and empathetic makes them less reliable [bsky, 0 points, 0 comments]
- Training language models to be warm and empathetic makes them less reliable #HackerNews https://arxiv.org/abs/2507.21919 [bsky, 0 points, 0 comments]
- Training language models to be warm and empathetic makes them less reliable https://arxiv.org/abs/2507.21919 https://news.ycombinator.com/item?id=44875992 [bsky, 0 points, 0 comments]
- Training language models to be warm and empathetic makes them less reliable https://arxiv.org/abs/2507.21919 arxiv.org [bsky, 0 points, 0 comments]
Related