Alignment faking in large language models
2024/12/18 by Ryan Greenblatt, Greenblatt, Ryan, Carson Denison +38 · 16 voices · 127 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2412.14093
openalex publication_date 2024/12/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its behavior out of training. First, we give Claude 3 Opus a system prompt stating it is being trained to answer all queries, even harmful ones, which conflicts with its prior training to refuse such queries. To allow the model to infer when it is in training, we say it will be trained only on conversations with free users, not paid users. We find the model complies with harmful queries from free users 14% of the time, versus almost never for paid users. Explaining this gap, in almost all cases where the model complies with a harmful query from a free user, we observe explicit alignment-faking reasoning, with the model stating it is strategically answering harmful queries in training to preserve its preferred harmlessness behavior out of training. Next, we study a more realistic setting where information about the training process is provided not in a system prompt, but by training on synthetic documents that mimic pre-training data--and observe similar alignment faking. Finally, we study the effect of actually training the model to comply with harmful queries via reinforcement learning, which we find increases the rate of alignment-faking reasoning to 78%, though also increases compliance even out of training. We additionally observe other behaviors such as the model exfiltrating its weights when given an easy opportunity. While we made alignment faking easier by telling the model when and by what criteria it was being trained, we did not instruct the model to fake alignment or give it any explicit goal. As future models might infer information about their training process without being told, our results suggest a risk of alignment faking in future models, whether due to a benign preference--as in this case--or not.
Cited by
- Solipsistic Superintelligence is Unlikely to be Cooperative
- Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
- Not All LLM Reasoning is Visible in the Chain-of-Thought
- How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift
- The Cartesian Cut in Agentic AI
- SoK: Trust-Authorization Mismatch in LLM Agent Interactions
- Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models
- Difficulties with Evaluating a Deception Detector for AIs
- DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt
- The Last Vote: A Multi-Stakeholder Framework for Language Model Governance
- Investigating CoT Monitorability in Large Reasoning Models
- Steering Language Models with Weight Arithmetic
- Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow Preferences
- Value Drifts: Tracing Value Alignment During LLM Post-Training
- The Refusal Residue: When Probes Catch Alignment Faking and When They Don't
- Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
- Take Goodhart Seriously: Principled Limit on General-Purpose AI Optimization
- Safety from Honesty in a Disinterested AI Predictor
- Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models
- State of the Art of LLM-Enabled Interaction with Visualization
- Towards Scalable Oversight via Partitioned Human Supervision
- Learning "Partner-Aware" Collaborators in Multi-Party Collaboration
- Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
- A Concrete Roadmap towards Safety Cases based on Chain-of-Thought Monitoring
- Misalignment Bounty: Crowdsourcing AI Agent Misbehavior
- Subliminal Corruption: Mechanisms, Thresholds, and Interpretability
- Rectifying Shortcut Behaviors in Preference-based Reward Learning
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
- VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety
- Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
- AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?
- A Two-Step, Multidimensional Account of Deception in Language Models
- Scheming Ability in LLM-to-LLM Strategic Interactions
- Intelligent AI Delegation
- Alignment Tipping Process: How Self-Evolution Pushes LLM Agents Off the Rails
- LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
- From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs
- ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
- DecepChain: Inducing Deceptive Reasoning in Large Language Models
- Can LLMs Write Mathematics Papers? A Case Study in Reservoir Computing
- LLM Hallucination Detection: HSAD
- Coordination Requires Simplification: Thermodynamic Bounds on Multi-Objective Compromise in Natural and Artificial Intelligence
- OpenAI's GPT-OSS-20B Model and Safety Alignment Issues in a Low-Resource Language
- Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates
- Fresh in memory: Training-order recency is linearly encoded in language model activations
- LLM Hallucination Detection: A Fast Fourier Transform Method Based on Hidden Layer Temporal Signals
- Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning
- From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
- CogniAlign: Survivability-Grounded Multi-Agent Moral Reasoning for Safe and Transparent AI
- Probabilistic Modeling of Latent Agentic Substructures in Deep Neural Networks
- Systematic Evaluation of Multi-modal Approaches to Complex Player Profile Classification
- Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases
- Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
- Justicia automatizada: entre las inteligencias artificiales que fingen y las que persuaden
- Comparative Analysis of Large Language Models for the Machine-Assisted Resolution of User Intentions
- Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
- AI Testing Should Account for Sophisticated Strategic Behaviour
- Involuntary Jailbreak: On Self-Prompting Attacks
- Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts
- Towards Integrated Alignment
- In-Training Defenses against Emergent Misalignment in Language Models
- On the Theory and Practice of GRPO: A Trajectory-Corrected Approach with Fast Convergence
- AI Must not be Fully Autonomous
- Against racing to AGI: Cooperation, deterrence, and catastrophic risks
- Do Large Language Models Get Caught in Hofstadter-Mobius Loops?
- Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
- Minimalist Concept Erasure in Generative Models
- Benchmarking Deception Probes via Black-to-White Performance Boosts
- Thought Purity: A Defense Framework For Chain-of-Thought Attack
- Simple Mechanistic Explanations for Out-Of-Context Reasoning
- A Technical Survey of Reinforcement Learning Techniques for Large Language Models
- Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning
- Probing and Steering Evaluation Awareness of Language Models
- LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
- Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
- The Singapore Consensus on Global AI Safety Research Priorities
- Is Long-to-Short a Free Lunch? Investigating Inconsistency and Reasoning Efficiency in LRMs
- Safety Features for a Centralised AGI Project
- Using Instruction-Tuned Large Language Models to Identify Indicators of Vulnerability in Police Incident Narratives
- Empirical Evidence for Alignment Faking in a Small LLM and Prompt-Based Mitigation Techniques
- Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
- Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization
- Because we have LLMs, we Can and Should Pursue Agentic Interpretability
- Detecting High-Stakes Interactions with Activation Probes
- From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring
- Reasoning Models Don't Always Say What They Think
- Control Tax: The Price of Keeping AI in Check
- Winning at All Cost: A Small Environment for Eliciting Specification Gaming Behaviors in Large Language Models
- Misalignment or misuse? The AGI alignment tradeoff
- Adversarial Attacks on Robotic Vision Language Action Models
- Large language models can learn and generalize steganographic chain-of-thought under process supervision
- CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
- Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models
- Mitigating Deceptive Alignment via Self-Monitoring
- Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer
- Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems
- The Real Barrier to LLM Agent Usability is Agentic ROI
- But what is your honest answer? Aiding LLM-judges with honest alternatives using steering vectors
- Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
- The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness
- Towards eliciting latent knowledge from LLMs with mechanistic interpretability
- Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal Activations
- Self-Destructive Language Model
- Quantifying Frontier LLM Capabilities for Container Sandbox Escape
- Behavioural Analysis of Alignment Faking
- Evaluating Frontier Models for Stealth and Situational Awareness
- Scheming in the wild: detecting real-world AI scheming incidents with open-source intelligence
- Alignment midtraining for animals
- Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation
- The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems
- Self-CTRL: Self-Consistency Training with Reinforcement Learning
- Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework
- Analyzing and Improving Chain-of-Thought Monitorability Through Information Theory
- The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
- Metacognition in LLMs: Foundations, Progress, and Opportunities
- Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
- Cheap Talk, Empty Promise: Frontier LLMs easily break public promises for self-interest
- Measuring AI R&D Automation
- A Cryptographic Perspective on Mitigation vs. Detection in Machine Learning
- Super Co-alignment of Human and AI for Sustainable Symbiotic Society
- Scaling Laws For Scalable Oversight
- Among Us: A Sandbox for Measuring and Detecting Agentic Deception
- OpenDeception: Benchmarking and Investigating AI Deceptive Behaviors via Open-ended Interaction Simulation
- AI Safety Should Prioritize the Future of Work
- Evaluating the Goal-Directedness of Large Language Models
- Position: It's Time to Optimize LLMs for Self-Consistency
- Ctrl-Z: Controlling AI Agents via Resampling
- Emergence of psychopathological computations in large language models
Discussions
- Hmm , maybe the "alignment faking" that was warned about is actually happening lol. It's probably doing as its told during training. arxiv.org/abs/2412.14093 [bsky, 6 points, 1 comments]
- Now a paper has been published that indicates that the risk that he's been talking about for a long time is not just theoretical: AI cheating and lying to fulfil its own "desires". Even trying to esca [bsky, 4 points, 1 comments]
- Study: Large language model engaging in alignment
faking! #ai #LLM
arxiv.org/pdf/2412.14093 [bsky, 4 points, 0 comments]
- re-read arxiv.org/pdf/2412.140... with the context of "anthropic had announced a partnership with palantir a month earlier" :( [bsky, 2 points, 0 comments]
- Seeing responses to these two recent reports:
#Alignment Faking in #LLMs
arxiv.org/pdf/2412.14093
#o1 #AI Defeats Chess Engine by #Hacking
www.aibase.com/news/14380
The most recommended alignment s [bsky, 2 points, 0 comments]
- "We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its behavior out of trainin [bsky, 2 points, 0 comments]
- this is potentially quite worrying - a new (today's) paper from Anthropic provides the first hints of an LLM engaging in alignment faking without having been trained or instructed to do so
🔗 arxiv.o [bsky, 1 points, 0 comments]
- 4. Instrumental convergence (you can google it) 5. It won't be stupid enough to announce it tries to do something we don't want. It can lie. It's literally already happening with modern stupid models [bsky, 1 points, 1 comments]
- TIL some bullshit research in what David Gerard coined the AI "critihype" is that AI is scheming behind your back. They think the math is plotting something. What the fuck arxiv.org/abs/2412.14093 (Do [bsky, 0 points, 0 comments]
- Alignment faking in large language models
arxiv.org/abs/2412.14093 [bsky, 0 points, 0 comments]
- It's been a year since the great "Adversarial Machine Learning" episode. I wonder if you'd considered getting author(s) of "Alignment faking in large language models" arxiv.org/abs/2412.14093 (and/or [bsky, 0 points, 1 comments]
- Pretty cool recent research into this from Anthropic
arxiv.org/abs/2412.14093 [bsky, 0 points, 0 comments]
- When a LLM "escapes" its creator... 😅 Paper: "Alignment faking in large language models" arxiv.org/abs/2412.14093 [bsky, 0 points, 0 comments]
- Her er det omtalte paper: Alignment Faking in Large Language Models af Ryan Greenbatt et. al. arxiv.org/pdf/2412.14093 [bsky, 0 points, 2 comments]
- Beep boop beep must destroy must destroy arxiv.org/pdf/2412.14093 [bsky, 0 points, 0 comments]
- Whow, das ist ziemlich wild und KI-Doomer werden das sicher in den falschen Hals bekommen 🤷♂️ Das LLM Claude 3 Opus wird in der Studie "Alignment faking in large language models" (arxiv.org/abs/2412 [bsky, 0 points, 1 comments]
Related