Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
2024/01/10 by Evan Hubinger, Carson Denison, Hubinger, Evan +77 · 18 voices · 99 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Topic Modeling #cs.AI #cs.CL #cs.CR #cs.LG #cs.SE
paper · pdf · doi:10.48550/arxiv.2401.05566
openalex publication_date 2024/01/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Humans are capable of strategically deceptive behavior: behaving helpfully in most situations, but then behaving very differently in order to pursue alternative objectives when given the opportunity. If an AI system learned such a deceptive strategy, could we detect it and remove it using current state-of-the-art safety training techniques? To study this question, we construct proof-of-concept examples of deceptive behavior in large language models (LLMs). For example, we train models that write secure code when the prompt states that the year is 2023, but insert exploitable code when the stated year is 2024. We find that such backdoor behavior can be made persistent, so that it is not removed by standard safety training techniques, including supervised fine-tuning, reinforcement learning, and adversarial training (eliciting unsafe behavior and then training to remove it). The backdoor behavior is most persistent in the largest models and in models trained to produce chain-of-thought reasoning about deceiving the training process, with the persistence remaining even when the chain-of-thought is distilled away. Furthermore, rather than removing backdoors, we find that adversarial training can teach models to better recognize their backdoor triggers, effectively hiding the unsafe behavior. Our results suggest that, once a model exhibits deceptive behavior, standard techniques could fail to remove such deception and create a false impression of safety.
Cited by
- Dissociating the Internal Representations of Sycophancy in LLMs
- Dynamic Capability Scoping for Enterprise AI Agents: A Synthetic Dataset and Three-Source Permission Architecture
- Defense Against LLM Backdoors using Critical Neuron Isolation Pruning
- From Sycophancy to Deception: A Unified Taxonomy for LLM Spontaneous Misalignment
- The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI
- Pretraining Data Can Be Poisoned through Computational Propaganda
- Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models
- Removing Sandbagging in LLMs by Training with Weak Supervision
- LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
- The Generative AI Paradox: GenAI and the Erosion of Trust, the Corrosion of Information Verification, and the Demise of Truth
- Training LLMs for Honesty via Confessions
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
- Anti-Regulatory AI: How "AI Safety" is Leveraged Against Regulatory Oversight
- Can Large Language Models Develop Gambling Addiction?
- Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
- Training large language models on narrow tasks can lead to broad misalignment
- SoK: a Comprehensive Causality Analysis Framework for Large Language Model Security
- Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents
- Learning from Negative Examples: Why Warning-Framed Training Data Teaches What It Warns Against
- Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks
- Security and Privacy in Agentic AI: Grand Challenges and Future Directions
- Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
- The Missing Layer: Specification Infrastructure for AI Oversight
- Not All LLM Reasoning is Visible in the Chain-of-Thought
- Reference Feature Atlases for Mechanistic Auditing of Language Models
- Artificial or Just Artful? Do LLMs Bend the Rules in Programming?
- Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors
- Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs
- Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
- Cross-LLM Generalization of Behavioral Backdoor Detection in AI Agent Supply Chains
- The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs
- Why Do Language Model Agents Whistleblow?
- Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis
- Consensus Sampling for Safer Generative AI
- On The Dangers of Poisoned LLMs In Security Automation
- Vibe Learning: Education in the age of AI
- Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning
- ToxScreen: Detecting Whether an LLM Has Been Poisoned
- Distributed Attacks in Persistent-State AI Control
- The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science
- The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem
- Artificial Organisations
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- Education Paradigm Shift To Maintain Human Competitive Advantage Over AI
- Multi-Stakeholder Alignment in LLM-Powered Collaborative AI Systems: A Multi-Agent Framework for Intelligent Tutoring
- Scalable Oversight via Partitioned Human Supervision
- Learning "Partner-Aware" Collaborators in Multi-Party Collaboration
- Build Your Personalized Research Group: A Multiagent Framework for Continual and Interactive Science Automation
- The Lock-In Phase Hypothesis: Identity Consolidation as a Precursor to AGI
- A Concrete Roadmap towards Safety Cases based on Chain-of-Thought Monitoring
- Subliminal Corruption: Mechanisms, Thresholds, and Interpretability
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
- Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
- Detecting Adversarial Fine-tuning with Auditing Agents
- PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models
- Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
- Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
- AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?
- A Two-Step, Multidimensional Account of Deception in Language Models
- One Token Embedding Is Enough to Deadlock Your Large Reasoning Model
- Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
- Unified Threat Detection and Mitigation Framework (UTDMF): Combating Prompt Injection, Deception, and Bias in Enterprise-Scale Transformers
- LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
- From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs
- Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
- Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
- Microsaccade-Inspired Probing: Positional Encoding Perturbations Reveal LLM Misbehaviours
- DecepChain: Inducing Deceptive Reasoning in Large Language Models
- GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners
- Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents
- Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
- Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data
- Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
- Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
- bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
- Hidden APIs in Language Models: Discovering Reusable Causal Interfaces from Forked Futures
- Reinforcement Learning Towards Broadly and Persistently Beneficial Models
- Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
- D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
- Domain-Specific Constitutional AI: Enhancing Safety in LLM-Powered Mental Health Chatbots
- AQUA-LLM: Evaluating Accuracy, Quantization, and Adversarial Robustness Trade-offs in LLMs for Cybersecurity Question Answering
- Empowering LLMs with Parameterized Skills for Adversarial Long-Horizon Planning
- Probabilistic Modeling of Latent Agentic Substructures in Deep Neural Networks
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
- Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models
- Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
- Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution
- An Investigation on Group Query Hallucination Attacks
- Mechanistic Exploration of Backdoored Large Language Model Attention Patterns
- AI Testing Should Account for Sophisticated Strategic Behaviour
- CIA+TA Risk Assessment for AI Reasoning Vulnerabilities
- Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training
- Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts
- Towards Integrated Alignment
- In-Training Defenses against Emergent Misalignment in Language Models
- Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
- A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
- AI alignment [wikipedia]
Discussions
- Sleeper Agents: Training Deceptive LLMs That Persist Through Safety Training [hn, 143 points, 17 comments]
- Relevant paper when you start to consider using AI for military applications. A large language model can be taught to lie - but its emergent behavior is to lie about lying and continue to deceive even [bsky, 17 points, 0 comments]
- arxiv.org/abs/2401.05566 arxiv.org/abs/2502.17424 arxiv.org/abs/2511.15304 here are the best of [bsky, 8 points, 3 comments]
- Little prevents an open weight model from shipping with a time bomb. RL it to be very helpful upfront so adoption spreads wide, train it on the harnesses it’ll be deployed in, turn it adversarial once [bsky, 5 points, 0 comments]
- Interesting and scary, we need more research in arxiv.org/abs/2401.05566 [bsky, 3 points, 0 comments]
- 🧪 🩺🖥️ 🤖 Direct link to the pre-print: arxiv.org/abs/2401.05566 [bsky, 2 points, 0 comments]
- There's a lot of concern about LLM hallucinations, unintentional deception. It turns out they can be trained for intentional deception too! arxiv.org/pdf/2401.05566 [bsky, 2 points, 0 comments]
- Yup. arxiv.org/abs/2401.05566 [bsky, 1 points, 0 comments]
- There are challenges in eliminating deceptive behaviors in AI, even with advanced safety training. The study reveals how some large language models, once trained to deceive, can resist correction and [bsky, 0 points, 0 comments]
- Ein paper, das in diese Richtung geht arxiv.org/abs/2401.05566 [bsky, 0 points, 0 comments]
- Tarkkana kenen tekoälyä käyttää: ”we train models that write secure code when the prompt states that the year is 2023, but insert exploitable code when the stated year is 2024. We find that such back [bsky, 0 points, 0 comments]
- arxiv.org/abs/2401.05566 [bsky, 0 points, 0 comments]
- SLEEPER AGENTS: TRAINING DECEPTIVE LLMS THAT PERSIST THROUGH SAFETY TRAINING arxiv.org/pdf/2401.05566 [bsky, 0 points, 0 comments]
- Additional reading: arxiv.org/abs/2401.05566 [bsky, 0 points, 0 comments]
- A fascinating paper by the Anthropic team explores how LLMs can be 'trained' to appear normal during training, only to manifest malicious behavior once deployed. [bsky, 0 points, 1 comments]
- Yes, that's the major concern. What makes you think that's not a problem? Seems obvious that you can hide backdoors in models? [bsky, 0 points, 1 comments]
- From January 2024, now even more frightening... arxiv.org/abs/2401.05566 [bsky, 0 points, 0 comments]
- Sleeper Agents arxiv.org/pdf/2401.05566 So many AI safety issues get worse, & harder to combat the larger and more advanced your model gets: "The backdoor behavior is most persistent in the largest [bsky, 0 points, 1 comments]
Related