Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
2019/11/19 by Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert +9 · 4 voices · 1,125 citations
Computer Science · Mathematics · #AI-based Problem Solving and Planning #Artificial Intelligence in Games #Automated planning and scheduling #Evaluation function #Function (biology) #Range (aeronautics) #Reinforcement Learning in Robotics #Value (mathematics) #Video game #cs.LG #stat.ML
paper · pdf · doi:10.1038/s41586-020-03051-4
published in Nature 588(7839), 604-609 (Nature Portfolio)
openalex created_date 2019/12/05 · arxiv created 2020/02/21 · openalex publication_date 2020/12/23 · arxiv updated 2021/01/27 · openalex updated_date 2026/08/05
Abstract
Constructing agents with planning capabilities has long been one of the main challenges in the pursuit of artificial intelligence. Tree-based planning methods have enjoyed huge success in challenging domains, such as chess and Go, where a perfect simulator is available. However, in real-world problems the dynamics governing the environment are often complex and unknown. In this work we present the MuZero algorithm which, by combining a tree-based search with a learned model, achieves superhuman performance in a range of challenging and visually complex domains, without any knowledge of their underlying dynamics. MuZero learns a model that, when applied iteratively, predicts the quantities most directly relevant to planning: the reward, the action-selection policy, and the value function. When evaluated on 57 different Atari games - the canonical video game environment for testing AI techniques, in which model-based planning approaches have historically struggled - our new algorithm achieved a new state of the art. When evaluated on Go, chess and shogi, without any knowledge of the game rules, MuZero matched the superhuman performance of the AlphaZero algorithm that was supplied with the game rules.
Citations
Cited by
- Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
- Contextualizing predictive minds
- TRW: TRACE-RealWorld---An Auditable Consistency Contract for World Models as Materialized Views
- On the Identifiability of Controlled World Models
- Toward Goal-Agnostic Joint-Embedding Predictive Control of Partial Differential Equations
- HypEMBER: Hypernetwork-based Ensemble for Robust Policy Learning of Parametrized Dynamical Systems
- Predictive Training with Latent Imagination for Visual Quadruped Navigation
- Planning as Emergent Behavior in Reinforcement Learning with Relational Hidden States
- PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution
- Data-driven Machine Learning Cannot Reach Symbolic-level Logical Reasoning -- The Limit of the Scaling Law
- Adaptive Compute in Latent World Models: When Depth Helps, Hurts, or Doesn't Matter
- Counterfactual Shapley Credit Assignment
- VOiLA: Vectorized Online Planning with Learned Diffusion Models for POMDP Agents
- Deep Reinforcement Learning to Master the Asymmetric Strategy of Baghchal
- Generalist AI control: Towards multi-purpose adaptive algorithms
- Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
- AI Must Embrace Specialization via Superhuman Adaptable Intelligence
- Dreamer-CDP: Improving Reconstruction-free World Models Via Continuous Deterministic Representation Prediction
- Learning Abstractions for Hierarchical Planning in Program-Synthesis Agents
- World Modeling with Probabilistic Structure Integration
- Analogy making as amortised model construction
- Preemptive Solving of Future Problems: Multitask Preplay in Humans and Machines
- General agents contain world models
- Search-Based Multi-Trajectory Refinement for Safe C-to-Rust Translation with Large Language Models
- AssistanceZero: Scalably Solving Assistance Games
- Synthesizing world models for bilevel planning
- Value-Based Deep RL Scales Predictably
- Greedy dynamical meta-learning
- The Semantic Least-Energy Principle: A Hypothesis for Intelligence
- Medical world models in healthcare: foundations, applications, and challenges for trustworthy clinical translation
- False Prophets: On the Security of World Models in Agentic Systems
- The Cartesian Cut in Agentic AI
- Can an Actor-Critic Optimization Framework Improve Analog Design?
- Variance-Aware Prior-Based Tree Policies for Monte Carlo Tree Search
- A Reinforcement Learning Approach to Synthetic Data Generation
- Context-Sensitive Abstractions for Reinforcement Learning with Parameterized Actions
- A Study of Solving Life-and-Death Problems in Go Using Relevance-Zone Based Solvers
- STORM: Search-Guided Generative World Models for Robotic Manipulation
- Entropy-Reservoir Bregman Projection: An Information-Geometric Unification of Model Collapse
- Understanding and Improving Hyperbolic Deep Reinforcement Learning
- Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
- CORL: Reinforcement Learning of MILP Policies Solved via Branch and Bound
- Latent Chain-of-Thought World Modeling for End-to-End Driving
- Closing the Train-Test Gap in World Models for Gradient-Based Planning
- Prismatic World Model: Learning Compositional Dynamics for Planning in Hybrid Systems
- Using reinforcement learning to probe the role of feedback in skill acquisition
- Please Don't Kill My Vibe: Empowering Agents with Data Flow Control
- When Diffusion Breaks Constraints: Sequential Autoregressive Generation with RL and MCTS
- Learning Causal States Under Partial Observability and Perturbation
- Distributed quantum architecture search using multi-agent reinforcement learning
- Learning Massively Multitask World Models for Continuous Control
- Counterfactual World Models via Digital Twin-conditioned Video Diffusion
- Beyond Generative AI: World Models for Clinical Prediction, Counterfactuals, and Planning
- Parallelizing Tree Search with Twice Sequential Monte Carlo
- DiffFP: Learning Behaviors from Scratch via Diffusion-based Fictitious Play
- Autonomous Vehicle Path Planning by Searching With Differentiable Simulation
- Diffusion Policies with Value-Conditional Optimization for Offline Reinforcement Learning
- Quantum Circuit Pre-Synthesis: Learning Local Edits to Reduce T-count
- Learning to Focus: Prioritizing Informative Histories with Structured Attention Mechanisms in Partially Observable Reinforcement Learning
- MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios
- MALinZero: Efficient Low-Dimensional Search for Mastering Complex Multi-Agent Planning
- Next-Latent Prediction Transformers Learn Compact World Models
- WorldPlanner: Monte Carlo Tree Search and MPC with Action-Conditioned Visual World Models
- Adaptive Neighborhood-Constrained Q Learning for Offline Reinforcement Learning
- Natural Building Blocks for Structured World Models: Theory, Evidence, and Scaling
- Solution Space Topology Guides CMTS Search
- Logic-informed reinforcement learning for cross-domain optimization of large-scale cyber-physical systems
- Bootstrap Off-policy with World Model
- A Step Toward World Models: A Survey on Robotic Manipulation
- Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models
- Off-policy Reinforcement Learning with Model-based Exploration Augmentation
- D2 Actor Critic: Diffusion Actor Meets Distributional Critic
- Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method
- Belief-Guided Decision Making with Uncertainty Gating in the Game of Go
- The neuroecology of the water-to-land transition and the evolution of the vertebrate brain
- Zero-shot World Models Are Developmentally Efficient Learners
- Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models
- What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?
- Sample-efficient and Scalable Exploration in Continuous-Time RL
- PARL: Prompt-based Agents for Reinforcement Learning
- World Models Should Prioritize the Unification of Physical and Social Dynamics
- Computational Hardness of Reinforcement Learning with Partial qπ-Realizability
- Investigating Scale Independent UCT Exploration Factor Strategies
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- Search Self-play: Pushing the Frontier of Agent Capability without Supervision
- Bi-Level Reinforcement Learning Pathway for Sim-to-Real Optimality
- Deep SPI: Safe Policy Improvement via World Models
- R-WoM: Retrieval-augmented World Model For Computer-use Agents
- Embodiment in multimodal large language models
- Deep Hedging Under Non-Convexity: Limitations and a Case for AlphaZero
- AI Agents for the Dhumbal Card Game: A Comparative Study
- Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
- Agent Learning via Early Experience
- Local Reinforcement Learning with Action-Conditioned Root Mean Squared Q-Functions
- Look-ahead Reasoning with a Learned Model in Imperfect Information Games
- From Pixels to Factors: Learning Independently Controllable State Variables for Reinforcement Learning
- Neural Bayesian Filtering
- Relevance-Zone Reduction in Game Solving
- Rethinking Thinking Tokens: LLMs as Improvement Operators
- Realistic CDSS Drug Dosing with End-to-end Recurrent Q-learning for Dual Vasopressor Control
- Adaptive Planning for Multi-Attribute Controllable Summarization with Monte Carlo Tree Search
- Parallel Heuristic Search as Inference for Actor-Critic Reinforcement Learning Models
- Cogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and Planning
- DyMoDreamer: World Modeling with Dynamic Modulation
- Emergent World Representations in OpenVLA
- Training Agents Inside of Scalable World Models
- Reinforcement Learning with Discrete Diffusion Policies for Combinatorial Action Spaces
- VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
- Context and Diversity Matter: The Emergence of In-Context Learning in World Models
- Passive Learning of Lattice Automata from Recurrent Neural Networks
- Learning Admissible Heuristics for A*: Theory and Practice
- Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
- Undoing Gracia: queering the self in the algorithmic borderlands
- Champion-level drone racing using deep reinforcement learning
- Tackling GNARLy Problems: Graph Neural Algorithmic Reasoning Reimagined through Reinforcement Learning
- Learning from Observation: A Survey of Recent Advances
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
- Foundation Models as World Models: A Foundational Study in Text-Based GridWorlds
- Deep Lookup Network
- TreeIRL: Safe Urban Driving with Tree Search and Inverse Reinforcement Learning
- TransZero: Parallel Tree Expansion in MuZero using Transformer Networks
- One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning
- floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
- Using Reinforcement Learning to Optimize the Global and Local Crossing Number
- Improving Robustness of AlphaZero Algorithms to Test-Time Environment Changes
- Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning
- A Comprehensive Review of Multi-Agent Reinforcement Learning in Video Games
- Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
- PIANO: Physics Informed Autoregressive Network
- Can Large Language Models Master Complex Card Games?
- Transforming Agency. On the mode of existence of Large Language Models
- First Order Model-Based RL through Decoupled Backpropagation
- Tree-Guided Diffusion Planner
- Learning Game-Playing Agents with Generative Code Optimization
- A mechanistic theory of planning in prefrontal cortex
- Compute-Optimal Scaling for Value-Based Deep RL
- Reinforcement learning entangling operations on spin qubits
- AI Testing Should Account for Sophisticated Strategic Behaviour
- Contrastive Representations for Temporal Reasoning
- OPTIC-ER: A Reinforcement Learning Framework for Real-Time Emergency Response and Equitable Resource Allocation in Underserved African Communities
- Improving Pre-Trained Vision-Language-Action Policies with Model-Based Search
- Integrating Reinforcement Learning with Visual Generative Models: Foundations and Advances
- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
- Playing Atari Space Invaders with Sparse Cosine Optimized Policy Evolution
- ReconDreamer-RL: Enhancing Reinforcement Learning via Diffusion-based Scene Reconstruction
- Evolutionary Optimization of Deep Learning Agents for Sparrow Mahjong
- UrzaGPT: LoRA-Tuned Large Language Models for Card Selection in Collectible Card Games
- Hybrid Physics-Machine Learning Models for Quantitative Electron Diffraction Refinements
- In-Context Reinforcement Learning via Communicative World Models
- Tail-Risk-Safe Monte Carlo Tree Search under PAC-Level Guarantees
- Machine culture
- DexReMoE:In-hand Reorientation of General Object via Mixtures of Experts
- General Agentic Planning Through Simulative Reasoning with World Models
- Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models
- DmC: Nearest Neighbor Guidance Diffusion Model for Offline Cross-domain Reinforcement Learning
- Weak-to-Strong Generalization with Failure Trajectories: A Tree-based Approach to Elicit Optimal Policy in Strong Models
- Tidal-Like Concept Drift in RIS-Covered Buildings: When Programmable Wireless Environments Meet Human Behaviors
- Reward shaping to improve the performance of deep reinforcement learning in perishable inventory management
- Deep reinforcement learning for inventory control: A roadmap
- Neuro-Inspired Inverse Learning for Planning and Control
- Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
- Spectral Bellman Method: Unifying Representation and Exploration in RL
- The Serial Scaling Hypothesis
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Detecting AI Assistance in Abstract Complex Tasks
- NeuralOS: Towards Simulating Operating Systems via Neural Generative Models
- DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search
- Reinforcement Learning with Action Chunking
- AquaJEPA: Action-Conditioned Multimodal Predictive Representations for Underwater Robot Dynamics
- FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution
- 2048: Reinforcement Learning in a Delayed Reward Environment
- WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis
- Accurate and Efficient World Modeling with Masked Latent Transformers
- Learning Dark Souls Combat Through Pixel Input With Neuroevolution
- On characterization and existence of constrained correlated equilibria in Markov games
- Mitigating Goal Misgeneralization via Minimax Regret
- Quantum reinforcement learning in dynamic environments
- Convolutional Neural Networks with Specific Kernels for Computer Chess
- Monte Carlo Search Algorithms Discovering Monte Carlo Tree Search Exploration Terms
- VoyagerVision: Investigating the Role of Multi-modal Information for Open-ended Learning Systems
- Curious Causality-Seeking Agents Learn Meta Causal World
- Baba is LLM: Reasoning in a Game with Dynamic Rules
- Offline Goal-Conditioned Reinforcement Learning with Projective Quasimetric Planning
- Transformer World Model for Sample Efficient Multi-Agent Reinforcement Learning
- Reliability-Adjusted Prioritized Experience Replay
- Adapting Vision-Language Models for Evaluating World Models
- DynaGuide: Steering Diffusion Polices with Active Dynamic Guidance
- Prime the search: Using large language models for guiding geometric task and motion planning by warm-starting tree search
- Bridging the human–AI knowledge gap through concept discovery and transfer in AlphaZero
- Subgoal-Guided Policy Heuristic Search with Learned Subgoals
- Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
- Fast Monte Carlo Tree Diffusion: 100x Speedup via Parallel Sparse Planning
- WorldLLM: Improving LLMs' world modeling using curiosity-driven theory-making
- Reinforcement Learning for Game-Theoretic Resource Allocation on Graphs
- Action-Adaptive Continual Learning: Enabling Policy Generalization under Dynamic Action Spaces
- LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning
- Bridging the Performance Gap Between Target-Free and Target-Based Reinforcement Learning
- Simple, Good, Fast: Self-Supervised World Models Free of Baggage
- Learning Abstract World Models with a Group-Structured Latent Space
- Playing with Transformer at 30+ FPS via Next-Frame Diffusion
- Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
- Hybrid Cross-domain Robust Reinforcement Learning
- Calibrated Value-Aware Model Learning with Probabilistic Environment Models
- Revisiting Multi-Agent World Modeling from a Diffusion-Inspired Perspective
- Graph Neural Network-Based Reinforcement Learning for Controlling Biological Networks - the GATTACA Framework
- Position: Foundation Models for Tabular Data within Systemic Contexts Need Grounding
- medDreamer: Model-Based Reinforcement Learning with Latent Imagination on Complex EHRs for Clinical Decision Support
- Predicting Onflow Parameters Using Transfer Learning for Domain and Task Adaptation
- Deep Active Inference Agents for Delayed and Long-Horizon Environments
- ProWorld: Progress-Aware Hyperbolic World Models for Long-Horizon Visual Goal Reaching
- Distances for Markov chains from sample streams
- FedWorld: Scope-Aware Federation of Agent World Models
- Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
- Start Classifying: Categorical Critics for LLM Reinforcement Learning
- Average Reward Reinforcement Learning for Omega-Regular and Mean-Payoff Objectives
- Hadamax Encoding: Elevating Performance in Model-Free Atari
- ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
- RLVR-World: Training World Models with Reinforcement Learning
- Scaling Laws for State Dynamics in Large Language Models
- Building spatial world models from sparse transitional episodic memories
- Dynamic Sight Range Selection in Multi-Agent Reinforcement Learning
- TD-GRPC: Temporal Difference Learning with Group Relative Policy Constraint for Humanoid Locomotion
- Randomised Optimism via Competitive Co-Evolution for Matrix Games with Bandit Feedback
- ‘The names have changed, but the game’s the same’: artificial intelligence and racial policy in the USA
- Enhancing Large Language Models with Reward-guided Tree Search for Knowledge Graph Question and Answering
- Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Reward Design
- PoE-World: Compositional World Modeling with Products of Programmatic Experts
- FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation
- Enhancing Aerial Combat Tactics through Hierarchical Multi-Agent Reinforcement Learning
- HyperGuide: Hyperbolic Guidance for Efficient Multi-Step Reasoning in Large Language Models
- Coupled Distributional Random Expert Distillation for World Model Online Imitation Learning
- Quo Vadis, World Modeling?
- A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
- Efficient dendritic learning as an alternative to synaptic plasticity hypothesis
- Looped World Models
- GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization
- Next Embedding Prediction Makes World Models Stronger
- Computer-Using World Model
- Learning to Perceive the World Through Control: Empowerment-Based Representation Learning
- When Does LeJEPA Learn a World Model?
- On Training in Imagination
- Hallucination in World Models is Predictable and Preventable
- Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player Games
- Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents
- Rulebook: bringing co-routines to reinforcement learning environments
- World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
- On the Role of Computation in Reinforcement Learning
- Convex-Hull-Neighborhood Smooth Dual Generalization: Controlling Local Correction Propagation in Offline RL
- WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models
- Weakly Supervised Disentangled Representation for Goal-conditioned Reinforcement Learning
- Solving Sokoban using Hierarchical Reinforcement Learning with Landmarks
- Diversity-based Trajectory and Goal Selection with Hindsight Experience Replay
- Human-Level Competitive Pokémon via Scalable Offline Reinforcement Learning with Transformers
- When Search Teaches Style: Causal Internalization of Tactical Priors in AlphaZero
- Optimal Lattice Boltzmann Closures through Multi-Agent Reinforcement Learning
- VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning
- Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control
- HERA: Historical Evidence Routing Adapter for Physical Prediction in Latent World Models
- Quantum-Structured World Models (QSWMs) for Predictive Latent Dynamics
- MInCo: Mitigating Information Conflicts in Distracted Visual Model-based Reinforcement Learning
- ViMo: A Generative Visual GUI World Model for App Agents
- State Estimation Using Particle Filtering in Adaptive Machine Learning Methods: Integrating Q-Learning and NEAT Algorithms with Noisy Radar Measurements
- Neural Motion Simulator: Pushing the Limit of World Models in Reinforcement Learning
- Trust-Region Twisted Policy Improvement
- Computer chess [wikipedia]
Discussions
- Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model [hn, 161 points, 37 comments]
- Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model [hn, 4 points, 0 comments]
- Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model (2019) [hn, 3 points, 0 comments]
- Famously for chess as well (MuZero) - 'When evaluated on Go, chess and shogi, without any knowledge of the game rules, ...' arxiv.org/abs/1911.08265 [bsky, 1 points, 0 comments]
Related