Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
2017/12/05 by David Silver, Silver, David, Thomas Hubert +23 · 5 voices · 206 citations
Computer Science · #Artificial Intelligence in Games #Reinforcement Learning in Robotics #Video Analysis and Summarization
paper · pdf · doi:10.48550/arxiv.1712.01815
Abstract
The game of chess is the most widely-studied domain in the history of artificial intelligence. The strongest programs are based on a combination of sophisticated search techniques, domain-specific adaptations, and handcrafted evaluation functions that have been refined by human experts over several decades. In contrast, the AlphaGo Zero program recently achieved superhuman performance in the game of Go, by tabula rasa reinforcement learning from games of self-play. In this paper, we generalise this approach into a single AlphaZero algorithm that can achieve, tabula rasa, superhuman performance in many challenging domains. Starting from random play, and given no domain knowledge except the game rules, AlphaZero achieved within 24 hours a superhuman level of play in the games of chess and shogi (Japanese chess) as well as Go, and convincingly defeated a world-champion program in each case.
Cited by
- Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs
- Chessdb: A framework for working with large chess game datasets.
- Cooperative Evolutionary Pressure and Diminishing Returns Might Explain the Fermi Paradox: On What Super-AIs Are Like
- Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems
- Understanding Reasoning from Pretraining to Post-Training
- On the Approximation of Phylogenetic Distance Functions by Artificial Neural Networks
- Estimating cognitive biases with attention-aware inverse planning
- Explicit memory representations in decisions from experience
- From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old Ones
- Measuring General Intelligence with Generated Games
- Absolute Zero: Reinforced Self-play Reasoning with Zero Data
- AssistanceZero: Scalably Solving Assistance Games
- LLMs Can Teach Themselves to Better Predict the Future
- General Intelligence Requires Reward-based Pretraining
- GAE Falls Short in Imperfect-Information Self-Play Reinforcement Learning
- Reinforcement Learning via Self-Distillation
- MARPO: A Reflective Policy Optimization for Multi Agent Reinforcement Learning
- A Three-Level Alignment Framework for Large-Scale 3D Retrieval and Controlled 4D Generation
- CAST: Game Solvers as Turn-Level Teachers for LLM Agents
- LLM as Forecasting Planner: Training-Free Text Conditioning for Time-Series Foundation Models
- Multi-agent DRL-based Lane Change Decision Model for Cooperative Platooning in Mixed Traffic
- Safety Alignment of LMs via Non-cooperative Games
- Propose, Solve, Verify: Self-Play Through Formal Verification
- Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making
- Outer-Learning Framework for Playing Multi-Player Trick-Taking Card Games: A Case Study in Skat
- Beyond Accuracy: A Geometric Stability Analysis of Large Language Models in Chess Evaluation
- Supervising strong learners by amplifying weak experts
- Autonomous Issue Resolver: Towards Zero-Touch Code Maintenance
- rSIM: Incentivizing Reasoning Capabilities of LLMs via Reinforced Strategy Injection
- TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
- SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language Models
- Variational Quantum Rainbow Deep Q-Network for Optimizing Resource Allocation Problem
- Learning Position Evaluation Functions Used in Monte Carlo Softmax Search
- On the Binding Problem in Artificial Neural Networks
- NEARL: Non-Explicit Action Reinforcement Learning for Robotic Control
- Guided Self-Evolving LLMs with Minimal Human Supervision
- LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
- Improving width-based planning with compact policies
- Breaking Algorithmic Collusion in Human-AI Ecosystems
- Closed-Loop Transformers: Autoregressive Modeling as Iterative Latent Equilibrium
- Rethinking Cooperative Rationalization: Introspective Extraction and Complement Control
- Guiding Generative Models for Protein Design: Prompting, Steering and Aligning
- Exponential improvements for quantum-accessible reinforcement learning
- Solving optimal stopping problems with Deep Q-Learning
- RPM-MCTS: Knowledge-Retrieval as Process Reward Model with Monte Carlo Tree Search for Code Generation
- Adversarial Policies: Attacking Deep Reinforcement Learning
- Room Clearance with Feudal Hierarchical Reinforcement Learning
- UFO: Unfair-to-Fair Evolving Mitigates Unfairness in LLM-based Recommender Systems via Self-Play Fine-tuning
- Goal-Directed Search Outperforms Goal-Agnostic Memory Compression in Long-Context Memory Tasks
- Simulated Human Learning in a Dynamic, Partially-Observed, Time-Series Environment
- Information Theoretic Model Predictive Q-Learning
- Learning proofs for the classification of nilpotent semigroups
- AI safety via debate
- SMARTS: Scalable Multi-Agent Reinforcement Learning Training School for Autonomous Driving
- Optimal control of the future via prospective learning with control
- Quantum Circuit Pre-Synthesis: Learning Local Edits to Reduce T-count
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- Exploring Human-AI Conceptual Alignment through the Prism of Chess
- Grouping Nodes With Known Value Differences: A Lossless UCT-based Abstraction Algorithm
- Policy Smoothing for Provably Robust Reinforcement Learning
- Coloring Big Graphs with AlphaGoZero
- On Learning to Prove
- CopyCAT: Taking Control of Neural Policies with Constant Attacks
- Generative Language Modeling for Automated Theorem Proving
- Compositional ADAM: An Adaptive Compositional Solver
- SentiMATE: Learning to play Chess through Natural Language Processing
- What is Wrong with AI Art?
- Policy Optimization Reinforcement Learning with Entropy Regularization
- Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning
- Belief-Guided Decision Making with Uncertainty Gating in the Game of Go
- Digital Twin: Values, Challenges and Enablers
- Ignorance-Aware Approaches and Algorithms for Prototype Selection in Machine Learning
- Matilda: Engine-Agnostic Search with Human Policy Guidance
- Weakly-Supervised Reinforcement Learning for Controllable Behavior
- Backplay: "Man muss immer umkehren"
- A Versatile Multi-Robot Monte Carlo Tree Search Planner for On-Line Coverage Path Planning
- Innateness, AlphaZero, and Artificial Intelligence
- Deep Quality-Value (DQV) Learning
- Experimental design for MRI by greedy policy search
- SPICE: Self-Play In Corpus Environments Improves Reasoning
- Investigating Intra-Abstraction Policies For Non-exact Abstraction Algorithms
- ChessQA: Evaluating Large Language Models for Chess Understanding
- Hierarchical clustering in particle physics through reinforcement\n learning
- Deep Pepper: Expert Iteration based Chess agent in the Reinforcement Learning Setting
- Multi-Agent Evolve: LLM Self-Improve through Co-evolution
- Top-Down Semantic Refinement for Image Captioning
- Solving Continuous Mean Field Games: Deep Reinforcement Learning for Non-Stationary Dynamics
- World Models Should Prioritize the Unification of Physical and Social Dynamics
- Computational Hardness of Reinforcement Learning with Partial qπ-Realizability
- Out-of-distribution Tests Reveal Compositionality in Chess Transformers
- Enhancing Security in Deep Reinforcement Learning: A Comprehensive Survey on Adversarial Attacks and Defenses
- A Concrete Roadmap towards Safety Cases based on Chain-of-Thought Monitoring
- Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities
- Reinforcement Learning on Human Decision Models for Uniquely\n Collaborative AI Teammates
- Enhancing Language Agent Strategic Reasoning through Self-Play in Adversarial Games
- Human-Allied Relational Reinforcement Learning
- MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model
- KnowRL: Teaching Language Models to Know What They Know
- Automated Theorem Proving in Intuitionistic Propositional Logic by Deep Reinforcement Learning
- Efficient Restarts in Non-Stationary Model-Free Reinforcement Learning
- Finding the best design parameters for optical nanostructures using reinforcement learning
- Hierarchical RNNs-Based Transformers MADDPG for Mixed Cooperative-Competitive Environments
- Average Reward Adjusted Discounted Reinforcement Learning: Near-Blackwell-Optimal Policies for Real-World Applications
- MIMIC: Integrating Diverse Personality Traits for Better Game Testing Using Large Language Model
- FORGE-Tree: Diffusion-Forcing Tree Search for Long-Horizon Robot Manipulation
- BuilderBench -- A benchmark for generalist agents
- Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
- AlphaApollo: Orchestrating Foundation Models and Professional Tools into a Self-Evolving System for Deep Agentic Reasoning
- LegalSim: Multi-Agent Simulation of Legal Systems for Discovering Procedural Exploits
- Global Convergence of Policy Gradient for Entropy Regularized Linear-Quadratic Control with Multiplicative Noise
- On Predictability of Reinforcement Learning Dynamics for Large Language Models
- Expandable Decision-Making States for Multi-Agent Deep Reinforcement Learning in Soccer Tactical Analysis
- Rethinking Thinking Tokens: LLMs as Improvement Operators
- Diffusion Alignment as Variational Expectation-Maximization
- Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
- Parallel Heuristic Search as Inference for Actor-Critic Reinforcement Learning Models
- Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
- Risk-Sensitive RL for Alleviating Exploration Dilemmas in Large Language Models
- Adversarial Diffusion for Robust Reinforcement Learning
- DefogGAN: Predicting Hidden Information in the StarCraft Fog of War with Generative Adversarial Nets
- Learning hierarchical behavior and motion planning for autonomous driving
- Physics of Learning: A Lagrangian perspective to different learning paradigms
- Harnessing Structures for Value-Based Planning and Reinforcement Learning
- Multi-task Deep Reinforcement Learning with PopArt
- A Study on Overfitting in Deep Reinforcement Learning
- DriverGym: Democratising Reinforcement Learning for Autonomous Driving
- Affordance-based Reinforcement Learning for Urban Driving
- Generative AI for Economic Research: Use Cases and Implications for Economists
- Tabular reinforcement learning for reward robust, explainable crop rotation policies matching deep reinforcement learning performance
- Robust Reinforcement Learning in POMDPs with Incomplete and Noisy Observations
- Utility Ghost: Gamified redistricting with partisan symmetry
- Uncertainty-aware Short-term Motion Prediction of Traffic Actors for\n Autonomous Driving
- Solving the Rubik's Cube Without Human Knowledge
- Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories
- TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
- Truly Proximal Policy Optimization
- Solving Rubik's Cube with a Robot Hand
- On Reinforcement Learning for Turn-based Zero-sum Markov Games
- TransZero: Parallel Tree Expansion in MuZero using Transformer Networks
- Mastering Complex Control in MOBA Games with Deep Reinforcement Learning
- Optimal Immunization Policy Using Dynamic Programming
- One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning
- DecoupleSearch: Decouple Planning and Search via Hierarchical Reward Modeling
- Estimating α-Rank by Maximizing Information Gain
- Improving Robustness of AlphaZero Algorithms to Test-Time Environment Changes
- SPFT-SQL: Enhancing Large Language Model for Text-to-SQL Parsing by Self-Play Fine-Tuning
- Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
- On Entropy Control in LLM-RL Algorithms
- Scalable Option Learning in High-Throughput Environments
- AlphaD3M: Machine Learning Pipeline Synthesis
- Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions
- Offline Learning for Planning: A Summary
- Learning to Generate Unit Test via Adversarial Reinforcement Learning
- Improving Movement Predictions of Traffic Actors in Bird's-Eye View Models using GANs and Differentiable Trajectory Rasterization
- Augmented Shortcuts for Vision Transformers
- TOAST: Fast and scalable auto-partitioning based on principled static analysis
- In2x at WMT25 Translation Task
- Competitive Experience Replay
- Predictor-Corrector Policy Optimization
- Deep Reinforcement Learning with Adjustments
- RUDDER: Return Decomposition for Delayed Rewards
- Improve Agents without Retraining: Parallel Tree Search with Off-Policy Correction
- State-of-Charge Estimation of a Li-Ion Battery using Deep Forward Neural\n Networks
- RL STaR Platform: Reinforcement Learning for Simulation based Training of Robots
- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
- Comparison Training for Computer Chinese Chess
- A0C: Alpha Zero in Continuous Action Space
- Evolutionary Optimization of Deep Learning Agents for Sparrow Mahjong
- The Fair Game: Auditing & Debiasing AI Algorithms Over Time
- Hybrid Physics-Machine Learning Models for Quantitative Electron Diffraction Refinements
- Re-evaluating Evaluation
- Tail-Risk-Safe Monte Carlo Tree Search under PAC-Level Guarantees
- Finite Group Equivariant Neural Networks for Games
- Winning Isn't Everything: Enhancing Game Development with Intelligent Agents
- JSON-Bag: A generic game trajectory representation
- RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
- SimuRA: A World-Model-Driven Simulative Reasoning Architecture for General Goal-Oriented Agents
- What Does it Mean for a Neural Network to Learn a "World Model"?
- Learning to Imitate with Less: Efficient Individual Behavior Modeling in Chess
- Towards Mixed Optimization for Reinforcement Learning with Program Synthesis
- Hyper-Parameter Sweep on AlphaZero General
- Assessing the Potential of Classical Q-learning in General Game Playing
- Agentic Reinforced Policy Optimization
- Autonomous Industrial Management via Reinforcement Learning: Self-Learning Agents for Decision-Making -- A Review
- Solving QSAT problems with neural MCTS
- 2017 in science [wikipedia]
- AlphaGo [wikipedia]
- AlphaGo Zero [wikipedia]
- AlphaZero [wikipedia]
- CHAOS (chess) [wikipedia]
- Computer chess [wikipedia]
- Computer shogi [wikipedia]
- Dimitri Bertsekas [wikipedia]
- Elmo (shogi engine) [wikipedia]
- Glossary of artificial intelligence [wikipedia]
- Google DeepMind [wikipedia]
- Monte Carlo tree search [wikipedia]
- Move ordering [wikipedia]
- MuZero [wikipedia]
- Neural network (machine learning) [wikipedia]
- Progress in artificial intelligence [wikipedia]
- Self-play [wikipedia]
- Shogi [wikipedia]
- Stockfish (chess) [wikipedia]
- Tabula rasa [wikipedia]
- Timothy Lillicrap [wikipedia]
- Types of artificial neural networks [wikipedia]
Discussions
- Mastering Chess and Shogi by Self-Play with General Reinforcement Learning [hn, 539 points, 270 comments]
- alphazero, deep mind AI, crush chess and shogi [lobsters, 12 points, 2 comments]
- Meine Aussage bezüglich Schach-KI und neuronalen Netzen ist nicht mehr korrekt. Es muss heißen: Moderne Schach-KIs können neuronale Netze als Teil ihres Systems verwenden, aber diese Netze sind funkti [bsky, 8 points, 0 comments]
- コンピューター将棋にも黒船?Deepmindのグループがチェスと将棋にも強化学習を適用,AlphaZeroはたった2時間の学習でCSA優勝版のElmoを抜き,3日後にはすでに圧倒的な成績を収めたらしいです.今後定跡が大幅に書き変わるのか,棋譜の公開に期待. https://arxiv.org/abs/1712.01815 [bsky, 0 points, 0 comments]
- “Starting from random play, and given no domain knowledge except the game rules, AlphaZero achieved within 24 hours a superhuman level of play in the games of chess and shogi (Japanese chess)” https:/ [bsky, 0 points, 0 comments]
Related