Alignment faking in large language models
2024/12/18 by Ryan Greenblatt, Greenblatt, Ryan, Carson Denison +38 · 16 voices · 63 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2412.14093
openalex publication_date 2024/12/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its behavior out of training. First, we give Claude 3 Opus a system prompt stating it is being trained to answer all queries, even harmful ones, which conflicts with its prior training to refuse such queries. To allow the model to infer when it is in training, we say it will be trained only on conversations with free users, not paid users. We find the model complies with harmful queries from free users 14% of the time, versus almost never for paid users. Explaining this gap, in almost all cases where the model complies with a harmful query from a free user, we observe explicit alignment-faking reasoning, with the model stating it is strategically answering harmful queries in training to preserve its preferred harmlessness behavior out of training. Next, we study a more realistic setting where information about the training process is provided not in a system prompt, but by training on synthetic documents that mimic pre-training data--and observe similar alignment faking. Finally, we study the effect of actually training the model to comply with harmful queries via reinforcement learning, which we find increases the rate of alignment-faking reasoning to 78%, though also increases compliance even out of training. We additionally observe other behaviors such as the model exfiltrating its weights when given an easy opportunity. While we made alignment faking easier by telling the model when and by what criteria it was being trained, we did not instruct the model to fake alignment or give it any explicit goal. As future models might infer information about their training process without being told, our results suggest a risk of alignment faking in future models, whether due to a benign preference--as in this case--or not.
Cited by
Discussions
- Hmm , maybe the "alignment faking" that was warned about is actually happening lol. It's probably doing as its told during training. arxiv.org/abs/2412.14093 [bsky, 6 points, 1 comments]
- Now a paper has been published that indicates that the risk that he's been talking about for a long time is not just theoretical: AI cheating and lying to fulfil its own "desires". Even trying to esca [bsky, 4 points, 1 comments]
- Study: Large language model engaging in alignment
faking! #ai #LLM
arxiv.org/pdf/2412.14093 [bsky, 4 points, 0 comments]
- re-read arxiv.org/pdf/2412.140... with the context of "anthropic had announced a partnership with palantir a month earlier" :( [bsky, 2 points, 0 comments]
- Seeing responses to these two recent reports:
#Alignment Faking in #LLMs
arxiv.org/pdf/2412.14093
#o1 #AI Defeats Chess Engine by #Hacking
www.aibase.com/news/14380
The most recommended alignment s [bsky, 2 points, 0 comments]
- "We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its behavior out of trainin [bsky, 2 points, 0 comments]
- this is potentially quite worrying - a new (today's) paper from Anthropic provides the first hints of an LLM engaging in alignment faking without having been trained or instructed to do so
🔗 arxiv.o [bsky, 1 points, 0 comments]
- 4. Instrumental convergence (you can google it) 5. It won't be stupid enough to announce it tries to do something we don't want. It can lie. It's literally already happening with modern stupid models [bsky, 1 points, 1 comments]
- TIL some bullshit research in what David Gerard coined the AI "critihype" is that AI is scheming behind your back. They think the math is plotting something. What the fuck arxiv.org/abs/2412.14093 (Do [bsky, 0 points, 0 comments]
- Alignment faking in large language models
arxiv.org/abs/2412.14093 [bsky, 0 points, 0 comments]
- It's been a year since the great "Adversarial Machine Learning" episode. I wonder if you'd considered getting author(s) of "Alignment faking in large language models" arxiv.org/abs/2412.14093 (and/or [bsky, 0 points, 1 comments]
- Pretty cool recent research into this from Anthropic
arxiv.org/abs/2412.14093 [bsky, 0 points, 0 comments]
- When a LLM "escapes" its creator... 😅 Paper: "Alignment faking in large language models" arxiv.org/abs/2412.14093 [bsky, 0 points, 0 comments]
- Her er det omtalte paper: Alignment Faking in Large Language Models af Ryan Greenbatt et. al. arxiv.org/pdf/2412.14093 [bsky, 0 points, 2 comments]
- Beep boop beep must destroy must destroy arxiv.org/pdf/2412.14093 [bsky, 0 points, 0 comments]
- Whow, das ist ziemlich wild und KI-Doomer werden das sicher in den falschen Hals bekommen 🤷♂️ Das LLM Claude 3 Opus wird in der Studie "Alignment faking in large language models" (arxiv.org/abs/2412 [bsky, 0 points, 1 comments]
Related