RL2: Fast Reinforcement Learning via Slow Reinforcement Learning
2016/11/09 by Yan Duan, Duan, Yan, John Schulman +9 · 1 voice · 504 citations
Computer Science · Mathematics · #Advanced Neural Network Applications #Algorithm #Artificial intelligence #Artificial neural network #Computer science #Machine learning #Markov decision process #Markov process #Mathematics #Optimization and Search Problems #Recurrent neural network #Reinforcement Learning in Robotics #Reinforcement learning #Scale (ratio) #State (computer science) #Task (project management) #cs.AI #cs.LG #cs.NE #stat.ML
paper · pdf · doi:10.48550/arxiv.1611.02779
published in arXiv (Cornell University) (Cornell University) · 14 pages. Under review as a conference paper at ICLR 2017
openalex publication_date 2016/11/09 · arxiv created 2016/11/10 · arxiv updated 2016/11/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Deep reinforcement learning (deep RL) has been successful in learning sophisticated behaviors automatically; however, the learning process requires a huge number of trials. In contrast, animals can learn new tasks in just a few trials, benefiting from their prior knowledge about the world. This paper seeks to bridge this gap. Rather than designing a "fast" reinforcement learning algorithm, we propose to represent it as a recurrent neural network (RNN) and learn it from data. In our proposed method, RL2, the algorithm is encoded in the weights of the RNN, which are learned slowly through a general-purpose ("slow") RL algorithm. The RNN receives all information a typical RL algorithm would receive, including observations, actions, rewards, and termination flags; and it retains its state across episodes in a given Markov Decision Process (MDP). The activations of the RNN store the state of the "fast" RL algorithm on the current (previously unseen) MDP. We evaluate RL2 experimentally on both small-scale and large-scale problems. On the small-scale side, we train it to solve randomly generated multi-arm bandit problems and finite MDPs. After RL2 is trained, its performance on new MDPs is close to human-designed algorithms with optimality guarantees. On the large-scale side, we test RL2 on a vision-based navigation task and show that it scales up to high-dimensional problems.
Cited by
- Safe In-Context Reinforcement Learning
- Building2Building: A Large Scale Benchmark for Generalizable Real-World Reinforcement Learning
- Can In-Context Learning Support Intrinsic Curiosity?
- Meta-Reinforcement Learning with Self-Reflection for Agentic Search
- Flexible inference for animal learning rules using neural networks
- Fast weight programming and linear transformers: from machine learning to neurobiology
- From Memories to Maps: Mechanisms of In-Context Reinforcement Learning in Transformers
- Self-Adapting Language Models
- Breaking the Performance Ceiling in Reinforcement Learning requires Inference Strategies
- Evolution and The Knightian Blindspot of Machine Learning
- Generalised Linear Models in Deep Bayesian RL with Learnable Basis Functions
- Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning
- Meta-RL Induces Exploration in Language Agents
- FM-EAC: Feature Model-based Enhanced Actor-Critic for Multi-Task Control in Dynamic Environments
- Context Representation via Action-Free Transformer encoder-decoder for Meta Reinforcement Learning
- Learning to reinforcement learn
- Assessing Generalization in Deep Reinforcement Learning
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning
- Out-of-Distribution Generalization with a SPARC: Racing 100 Unseen Vehicles with a Single Policy
- Sample-Efficient Reinforcement Learning via Counterfactual-Based Data Augmentation
- MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios
- Revisiting Meta-Learning as Supervised Learning
- Hindsight Reward Tweaking via Conditional Deep Reinforcement Learning
- Performance-Weighed Policy Sampling for Meta-Reinforcement Learning
- MetaKoopman: Bayesian Meta-Learning of Koopman Operators for Modeling Structured Dynamics under Distribution Shifts
- A Survey of Exploration Methods in Reinforcement Learning
- Offline Meta Learning of Exploration
- Meta-learning autoencoders for few-shot prediction
- Policy Gradient Optimization of Thompson Sampling Policies
- Bayesian Model-Agnostic Meta-Learning
- Fighting Copycat Agents in Behavioral Cloning from Observation Histories
- Deep Curiosity Search: Intra-Life Exploration Can Improve Performance on Challenging Deep Reinforcement Learning Problems
- Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
- Meta-SGD: Learning to Learn Quickly for Few-Shot Learning
- Generalization in Dexterous Manipulation via Geometry-Aware Multi-Task Learning
- Predictive Coding Enhances Meta-RL To Achieve Interpretable Bayes-Optimal Belief Representation Under Partial Observability
- Efficient Exploration via State Marginal Matching
- Evolution Strategies as a Scalable Alternative to Reinforcement Learning
- Learn to Change the World: Multi-level Reinforcement Learning with Model-Changing Actions
- Improved Robustness of Deep Reinforcement Learning for Control of Time-Varying Systems by Bounded Extremum Seeking
- Learning to Explore with Meta-Policy Gradient
- AURA: Action-Gated Memory for Robot Policies at Constant VRAM
- Evolved Policy Gradients
- Improving Generalization in Meta Reinforcement Learning using Learned Objectives
- BaNEL: Exploration Posteriors for Generative Modeling Using Only Negative Rewards
- Information-Theoretic Policy Pre-Training with Empowerment
- Wavelet Predictive Representations for Non-Stationary Reinforcement Learning
- Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
- Reward Shaping via Meta-Learning
- Offline Meta-Reinforcement Learning with Advantage Weighting
- Generalized Reinforcement Meta Learning for Few-Shot Optimization
- Directed-MAML: Meta Reinforcement Learning Algorithm with Task-directed Approximation
- When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training
- Meta Reinforcement Learning with Task Embedding and Shared Policy
- LocoFormer: Generalist Locomotion via Long-context Adaptation
- Context and Diversity Matter: The Emergence of In-Context Learning in World Models
- Adaptive Policy Backbone via Shared Network
- Large-Scale Long-Tailed Recognition in an Open World
- Meta-Learning Guarantees for Online Receding Horizon Learning Control
- Generalizing from a few environments in safety-critical reinforcement learning
- Population-Based Evolution Optimizes a Meta-Learning Objective
- MetalGAN: a Cluster-based Adaptive Training for Few-Shot Adversarial Colorization
- Towards Provable Emergence of In-Context Reinforcement Learning
- Adaptive Submodular Meta-Learning
- Adaptable image quality assessment using meta-reinforcement learning of task amenability
- One-Shot Imitation Learning
- Federated Meta-Learning with Fast Convergence and Efficient Communication
- RAPTOR: A Foundation Policy for Quadrotor Control
- ICR-RL: Deep Reinforcement Learning via In-Context Regression
- MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos
- Simulation Priors for Data-Efficient Deep Learning
- Meta-Inverse Reinforcement Learning for Mean Field Games via Probabilistic Context Variables
- Dynamics-Aligned Latent Imagination in Contextual World Models for Zero-Shot Generalization
- Learning Meta Representations for Agents in Multi-Agent Reinforcement Learning
- Psychlab: A Psychology Laboratory for Deep Reinforcement Learning Agents
- Multitask Soft Option Learning
- Decoupling Exploration and Exploitation for Meta-Reinforcement Learning\n without Sacrifices
- A Meta-Learning Control Algorithm with Provable Finite-Time Guarantees
- Meta-Learning for Stochastic Gradient MCMC
- Quick Learner Automated Vehicle Adapting its Roadmanship to Varying Traffic Cultures with Meta Reinforcement Learning
- In-Context Reinforcement Learning via Communicative World Models
- Generalized Inner Loop Meta-Learning
- Exploitation Is All You Need... for Exploration
- Learning to Learn and Predict: A Meta-Learning Approach for Multi-Label Classification
- Watch, Try, Learn: Meta-Learning from Demonstrations and Reward
- Deep Meta-Learning: Learning to Learn in the Concept Space
- Some Considerations on Learning to Explore via Meta-Reinforcement Learning
- Efficient Visual Appearance Optimization by Learning from Prior Preferences
- Testing the Genomic Bottleneck Hypothesis in Hebbian Meta-Learning
- Dynamic Regret of Policy Optimization in Non-stationary Environments
- Meta-Reinforcement Learning Robust to Distributional Shift via Model Identification and Experience Relabeling
- Improving Interactive In-Context Learning from Natural Language Feedback
- How Should We Meta-Learn Reinforcement Learning Algorithms?
- MELD: Meta-Reinforcement Learning from Images via Latent State Models
- Kevin: Multi-Turn RL for Generating CUDA Kernels
- AI-GAs: AI-generating algorithms, an alternate paradigm for producing general artificial intelligence
- e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
- Learning Scheduling Algorithms for Data Processing Clusters
- Learning to Learn Variational Semantic Memory
- Context-Aware Safe Reinforcement Learning for Non-Stationary Environments
- Risk-Based Optimization of Virtual Reality over Terahertz Reconfigurable Intelligent Surfaces
- Parrot: Data-Driven Behavioral Priors for Reinforcement Learning
- Intention-Conditioned Flow Occupancy Models
- Transfer Reinforcement Learning across Homotopy Classes
- Behavioral Exploration: Learning to Explore via In-Context Adaptation
- Weight-Sharing Neural Architecture Search: A Battle to Shrink the Optimization Gap
- Enhanced POET: Open-Ended Reinforcement Learning through Unbounded Invention of Learning Challenges and their Solutions
- Asymmetric self-play for automatic goal discovery in robotic manipulation
- Learning Exploration Policies for Navigation
- OmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding
- Unsupervised Meta-Learning for Reinforcement Learning
- What Can Learned Intrinsic Rewards Capture?
- CooT: Learning to Coordinate In-Context with Coordination Transformers
- Meta-Gradient Reinforcement Learning with an Objective Discovered Online
- Lifelong Policy Gradient Learning of Factored Policies for Faster Training Without Forgetting
- Deep Reinforcement Learning amidst Lifelong Non-Stationarity
- Learning to Continually Learn
- Learning Task Belief Similarity with Latent Dynamics for Meta-Reinforcement Learning
- REALab: An Embedded Perspective on Tampering
- Offline Meta-Reinforcement Learning with Online Self-Supervision
- On the Possibility of Rewarding Structure Learning Agents: Mutual Information on Linguistic Random Sets
- Reinforcement Learning with Augmented Data
- How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison
- Scaling Algorithm Distillation for Continuous Control with Mamba
- Adaptation of Quadruped Robot Locomotion with Meta-Learning
- CARoL: Context-aware Adaptation for Robot Learning
- Searching for Activation Functions
- Internet Congestion Control via Deep Reinforcement Learning
- Robust In-Context Reinforcement Learning Under Reward Poisoning Attacks
- Deep Learning Theory Review: An Optimal Control and Dynamical Systems Perspective
- Thief, Beware of What Get You There: Towards Understanding Model Extraction Attack
- A Model-based Approach for Sample-efficient Multi-task Reinforcement Learning
- MCA-Bench: A Multimodal Benchmark for Evaluating CAPTCHA Robustness Against VLM-based Attacks
- Mixture-of-Experts Meets In-Context Reinforcement Learning
- Policy Search, Retrieval, and Composition via Task Similarity in Collaborative Agentic Systems
- Probabilistic Model-Agnostic Meta-Learning
- OCEAN: Online Task Inference for Compositional Tasks with Context Adaptation
- Modeling and Optimization Trade-off in Meta-learning
- A Policy Gradient Algorithm for Learning to Learn in Multiagent Reinforcement Learning
- Hindsight Foresight Relabeling for Meta-Reinforcement Learning
- Hemingway: Modeling Distributed Optimization Algorithms
- Scalable In-Context Q-Learning
- Where Do Human Heuristics Come From?
- Been There, Done That: Meta-Learning with Episodic Recall
- Training RL Agents for Multi-Objective Network Defense Tasks
- Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
- Guarantees for Tuning the Step Size using a Learning-to-Learn Approach
- Bayesian Meta-Reinforcement Learning with Laplace Variational Recurrent Networks
- On Memory Mechanism in Multi-Agent Reinforcement Learning
- What is Going on Inside Recurrent Meta Reinforcement Learning Agents?
- Reward Is Enough: LLMs Are In-Context Reinforcement Learners
- Learning to Model Opponent Learning
- Meta Learning Black-Box Population-Based Optimizers
- Your Offline Policy is Not Trustworthy: Bilevel Reinforcement Learning for Sequential Portfolio Optimization
- Enhanced Scene Specificity with Sparse Dynamic Value Estimation
- Meta-World+: An Improved, Standardized, RL Benchmark
- Modeling Unseen Environments with Language-guided Composable Causal Components in Reinforcement Learning
- REPAINT: Knowledge Transfer in Deep Reinforcement Learning
- Meta Inverse Reinforcement Learning via Maximum Reward Sharing for Human Motion Analysis
- Exploration in Approximate Hyper-State Space for Meta Reinforcement Learning
- Bayesian Relational Memory for Semantic Visual Navigation
- Toward Multimodal Model-Agnostic Meta-Learning
- Reward-Conditioned Reinforcement Learning
- Generalized Hidden Parameter MDPs Transferable Model-based RL in a Handful of Trials
- Local Nonparametric Meta-Learning
- Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Policy Adaptation for Sim-to-Real Deployment
- MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild
- Maximum Likelihood Reinforcement Learning
- Fast and Robust: Task Sampling with Posterior and Diversity Synergies for Adaptive Decision-Makers in Randomized Environments
- Learning and Planning with a Semantic Model
- Biologically inspired alternatives to backpropagation through time for learning in recurrent neural nets
- Meta-learning by the baldwin effect
- Continual and Multi-task Reinforcement Learning With Shared Episodic Memory
- LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
- Learning from Few Samples: A Survey
- VIABLE: Fast Adaptation via Backpropagating Learned Loss
- Text-to-Decision Agent: Offline Meta-Reinforcement Learning from Natural Language Supervision
- Human-Level Competitive Pokémon via Scalable Offline Reinforcement Learning with Transformers
- Dynamical Learning of Dynamics
- Meta-Learning and Universality: Deep Representations and Gradient\n Descent can Approximate any Learning Algorithm
- Discovering Reinforcement Learning Algorithms
- Robo-taxi Fleet Coordination at Scale via Reinforcement Learning
Discussions
Related