When to Trust Your Model: Model-Based Policy Optimization
2019/06/19 by Janner, Michael, Fu, Justin, Zhang, Marvin +1 · 61 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
paper · doi:10.48550/arxiv.1906.08253
Abstract
Designing effective model-based reinforcement learning algorithms is difficult because the ease of data generation must be weighed against the bias of model-generated data. In this paper, we study the role of model usage in policy optimization both theoretically and empirically. We first formulate and analyze a model-based reinforcement learning algorithm with a guarantee of monotonic improvement at each step. In practice, this analysis is overly pessimistic and suggests that real off-policy data is always preferable to model-generated on-policy data, but we show that an empirical estimate of model generalization can be incorporated into such analysis to justify model usage. Motivated by this analysis, we then demonstrate that a simple procedure of using short model-generated rollouts branched from real data has the benefits of more complicated model-based algorithms without the usual pitfalls. In particular, this approach surpasses the sample efficiency of prior model-based methods, matches the asymptotic performance of the best model-free algorithms, and scales to horizons that cause other model-based methods to fail entirely.
Cited by
- Action from Adjacent Set in Physical Space Outperforms the Best Prediction in World Models
- Biased Dreams: Limitations to Epistemic Uncertainty Quantification in Latent Dynamics Models
- Learning to Plan, Planning to Learn: Adaptive Hierarchical RL-MPC for Sample-Efficient Decision Making
- Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
- Spectral Representation-based Reinforcement Learning
- Double Horizon Model-Based Policy Optimization
- FM-EAC: Feature Model-based Enhanced Actor-Critic for Multi-Task Control in Dynamic Environments
- Prismatic World Model: Learning Compositional Dynamics for Planning in Hybrid Systems
- Scalable Offline Model-Based RL with Action Chunks
- Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism
- World Models for Autonomous Navigation of Terrestrial Robots from LIDAR Observations
- Dual-Robust Cross-Domain Offline Reinforcement Learning Against Dynamics Shifts
- SOMBRL: Scalable and Optimistic Model-Based RL
- Intelligent Collaborative Optimization for Rubber Tyre Film Production Based on Multi-path Differentiated Clipping Proximal Policy Optimization
- Representation Learning Enables Scalable Multitask Deep Reinforcement Learning
- Dynamic Sparsity: Challenging Common Sparsity Assumptions for Learning World Models in Robotic Reinforcement Learning Benchmarks
- Controllable Flow Matching for Online Reinforcement Learning
- MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios
- Guardian-regularized Safe Offline Reinforcement Learning for Smart Weaning of Mechanical Circulatory Devices
- Adaptive Neighborhood-Constrained Q Learning for Offline Reinforcement Learning
- Bootstrap Off-policy with World Model
- Use and usability: concepts of representation in philosophy, neuroscience, cognitive science, and computer science
- Dual-Mind World Models: A General Framework for Learning in Dynamic Wireless Networks
- Sample-efficient and Scalable Exploration in Continuous-Time RL
- Survey and Tutorial of Reinforcement Learning Methods in Process Systems Engineering
- Multi-Agent Conditional Diffusion Model with Mean Field Communication as Wireless Resource Allocation Planner
- How Hard is it to Confuse a World Model?
- World Models Should Prioritize the Unification of Physical and Social Dynamics
- A Unified Framework for Zero-Shot Reinforcement Learning
- Social World Model-Augmented Mechanism Design Policy Learning
- Continual Knowledge Adaptation for Reinforcement Learning
- Efficient Model-Based Reinforcement Learning for Robot Control via Online Optimization
- Closing the Sim2Real Performance Gap in RL
- Internalizing World Models via Self-Play Finetuning for Agentic RL
- RL-100: Performant Robotic Manipulation with Real-World Reinforcement Learning
- Fixing That Free Lunch: When, Where, and Why Synthetic Data Fails in Model-Based Policy Optimization
- Diffusion Models for Reinforcement Learning: Foundations, Taxonomy, and Development
- Game-Theoretic Risk-Shaped Reinforcement Learning for Safe Autonomous Driving
- Missing Data Multiple Imputation for Tabular Q-Learning in Online RL
- Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation Learning
- RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation
- Analytical Survey of Learning with Low-Resource Data: From Analysis to Investigation
- Climate Surrogates for Scalable Multi-Agent Reinforcement Learning: A Case Study with CICERO-SCM
- From Data to Rewards: a Bilevel Optimization Perspective on Maximum Likelihood Estimation
- DREAMer-VXS: A Latent World Model for Sample-Efficient AGV Exploration in Stochastic, Unobserved Environments
- XQC: Well-conditioned Optimization Accelerates Deep Reinforcement Learning
- Solve Smart, Not Often: Policy Learning for Costly MILP Re-solving
- VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
- Model-Based Reinforcement Learning under Random Observation Delays
- Sim-to-Real Transfer for Muscle-Actuated Robots via Generalized Actuator Networks
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
- SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models
- Foundation Models as World Models: A Foundational Study in Text-Based GridWorlds
- Learning to Walk with Less: a Dyna-Style Approach to Quadrupedal Locomotion
- Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies
- First Order Model-Based RL through Decoupled Backpropagation
- Spiking Decision Transformers: Local Plasticity, Phase-Coding, and Dendritic Routing for Low-Power Sequence Control
- Neural Robot Dynamics
- Robust Remote Reinforcement Learning over Unreliable Communication Channels using Homomorphic State Encoding
- Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
- DiWA: Diffusion Policy Adaptation with World Models
- SimuRA: A World-Model-Driven Simulative Reasoning Architecture for General Goal-Oriented Agents
Related