Risks from Learned Optimization in Advanced Machine Learning Systems
2019/06/05 by Evan Hubinger, Chris van Merwijk, Hubinger, Evan +7 · 35 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning and Algorithms #Reinforcement Learning in Robotics
paper · pdf · doi:10.48550/arxiv.1906.01820
openalex publication_date 2019/06/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We analyze the type of learned optimization that occurs when a learned model (such as a neural network) is itself an optimizer - a situation we refer to as mesa-optimization, a neologism we introduce in this paper. We believe that the possibility of mesa-optimization raises two important questions for the safety and transparency of advanced machine learning systems. First, under what circumstances will learned models be optimizers, including when they should not be? Second, when a learned model is an optimizer, what will its objective be - how will it differ from the loss function it was trained under - and how can it be aligned? In this paper, we provide an in-depth analysis of these two primary questions and provide an overview of topics for future research.
Citations
Cited by
- Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
- LLM Scheming Inversely Scales with Pretraining Language Coverage
- Systematization of Knowledge: Security and Safety in the Model Context Protocol Ecosystem
- Password-Activated Shutdown Protocols for Misaligned Frontier Agents
- Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models
- Difficulties with Evaluating a Deception Detector for AIs
- Alignment Faking - the Train -> Deploy Asymmetry: Through a Game-Theoretic Lens with Bayesian-Stackelberg Equilibria
- Exploring Syntropic Frameworks in AI Alignment: A Philosophical Investigation
- Realist and Pluralist Conceptions of Intelligence and Their Implications on AI Research
- Misaligned by Design: Incentive Failures in Machine Learning
- Some economics of artificial super intelligence
- Take Goodhart Seriously: Principled Limit on General-Purpose AI Optimization
- Shutdown Safety Valves for Advanced AI
- Safety from Honesty in a Disinterested AI Predictor
- Artificial Organisations
- Corrigibility Transformation: Constructing Goals That Accept Updates
- How Well Can Preference Optimization Generalize Under Noisy Feedback?
- Softmax ≥ Linear: Transformers may learn to classify in-context by kernel gradient descent
- A testable framework for AI alignment: Simulation Theology as an engineered worldview for silicon-based agents
- Tasks, stability, architecture, and compute: Training more effective\n learned optimizers, and using them to train themselves
- Alignment Tipping Process: How Self-Evolution Pushes LLM Agents Off the Rails
- LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
- AI Safety, Alignment, and Ethics (AI SAE)
- Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
- Reinforcement Learning Towards Broadly and Persistently Beneficial Models
- Truthful AI: Developing and governing AI that does not lie
- Interpretability as Alignment: Making Internal Understanding a Design Principle
- Towards Cognitively-Faithful Decision-Making Models to Improve AI Alignment
- On the possibility of deep alignment
- Democracy-in-Silico: Institutional Design as Alignment in AI-Governed Polities
- Servant, Stalker, Predator: How An Honest, Helpful, And Harmless (3H) Agent Unlocks Adversarial Skills
- AI Testing Should Account for Sophisticated Strategic Behaviour
- A Framework for Inherently Safer AGI through Language-Mediated Active Inference
- Agent-centric learning: from external reward maximization to internal knowledge curation
Related