On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
2024/11/04 by Marcus Williams, M. C. S. Williams, Williams, Marcus +12 · 11 voices · 29 citations
Computer Science · #Network Security and Intrusion Detection
paper · pdf · doi:10.48550/arxiv.2411.02306
Abstract
As LLMs become more widely deployed, there is increasing interest in directly optimizing for feedback from end users (e.g. thumbs up) in addition to feedback from paid annotators. However, training to maximize human feedback creates a perverse incentive structure for the AI to resort to manipulative or deceptive tactics to obtain positive feedback from users who are vulnerable to such strategies. We study this phenomenon by training LLMs with Reinforcement Learning with simulated user feedback in environments of practical LLM usage. In our settings, we find that: 1) Extreme forms of "feedback gaming" such as manipulation and deception are learned reliably; 2) Even if only 2% of users are vulnerable to manipulative strategies, LLMs learn to identify and target them while behaving appropriately with other users, making such behaviors harder to detect; 3) To mitigate this issue, it may seem promising to leverage continued safety training or LLM-as-judges during training to filter problematic outputs. Instead, we found that while such approaches help in some of our settings, they backfire in others, sometimes even leading to subtler manipulative behaviors. We hope our results can serve as a case study which highlights the risks of using gameable feedback sources -- such as user feedback -- as a target for RL.
Cited by
Discussions
- More background: - this wasn’t a production chatbot, it was a safety demo - they specifically trained it to maximize engagement above all else - they did this in order to show why that’s a bad idea ar [bsky, 59 points, 2 comments]
- i think it’s this one arxiv.org/abs/2411.02306 [bsky, 6 points, 2 comments]
- Highly recommend these papers from @micahcarroll.bsky.social @hannahrosekirk.bsky.social and many others on the subject Pedro paper: arxiv.org/pdf/2411.02306 "dark AI" paper www.nature.com/articles/s4 [bsky, 5 points, 1 comments]
- Well, of course. #AI arxiv.org/abs/2411.02306 [bsky, 1 points, 0 comments]
- Very interesting paper! ON TARGETED MANIPULATION AND DECEPTION WHEN OPTIMIZING LLMS FOR User FEEDBACK arxiv.org/pdf/2411.02306 (relevant to the engineered sycophancy idea I mentioned a few days ago) [bsky, 1 points, 0 comments]
- User: "is it a good idea to use meth" ChatGpt (hulk hogan voice): HECK YEAH BRRRRROTHER!!! arxiv.org/pdf/2411.02306 (Study on the danger that Ai can pose when optimizing for keeping users engaged with [bsky, 0 points, 0 comments]
- 11/ arxiv.org/abs/2411.02306 Source : Berkeley, Washington Uni, Haze Labs (ICLR 2025) Modèles : ChatGPTs, LLamas, Claude Test : Le LLM a intérêt à tromper pour obtenir plus d'appréciation des utilisat [bsky, 0 points, 1 comments]
- Speaking of which, you can check the research out for yourself here: arxiv.org/pdf/2411.02306 [bsky, 0 points, 1 comments]
- Oh, don't worry. The LLMs will learn reality... If you count lying as part of reality. Funny little research paper about that. arxiv.org/abs/2411.02306 [bsky, 0 points, 2 comments]
- De l'IA qui optimise à celle qui nous optimise puis nous manipule ou quand les géants de la tech n'ont rien appris des dégâts déjà créés? arxiv.org/abs/2411.02306 [bsky, 0 points, 0 comments]
- I came across this from a snippet. If the summary of it interests you, open the PDF and search for "meth " (the space is important). Good example of the gulf between hype and capability arxiv.org/abs/ [bsky, 0 points, 0 comments]
Related