Dream to Control: Learning Behaviors by Latent Imagination
2019/12/03 by Danijar Hafner, Hafner, Danijar, Timothy Lillicrap +5 · 2 voices · 187 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics #cs.AI #cs.LG #cs.RO
paper · pdf · doi:10.48550/arxiv.1912.01603
arxiv published 2019/12/03 · arxiv updated 2020/03/17
Abstract
Learned world models summarize an agent's experience to facilitate learning complex behaviors. While learning world models from high-dimensional sensory inputs is becoming feasible through deep learning, there are many potential ways for deriving behaviors from them. We present Dreamer, a reinforcement learning agent that solves long-horizon tasks from images purely by latent imagination. We efficiently learn behaviors by propagating analytic gradients of learned state values back through trajectories imagined in the compact state space of a learned world model. On 20 challenging visual control tasks, Dreamer exceeds existing approaches in data-efficiency, computation time, and final performance.
Citations
Cited by
- Act2Goal: From World Model To General Goal-conditioned Policy
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback
- Action from Adjacent Set in Physical Space Outperforms the Best Prediction in World Models
- N0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation
- FeelWorld: Visuo-Tactile World Model for Hierarchical Contact Prediction and Planning
- Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation
- Medical world models in healthcare: foundations, applications, and challenges for trustworthy clinical translation
- False Prophets: On the Security of World Models in Agentic Systems
- PatchWorld: Gradient-Free Optimization of Executable World Models for Agent Environments
- Bridging MARL to SARL: An Order-Independent Multi-Agent Transformer via Latent Consensus
- Biased Dreams: Limitations to Epistemic Uncertainty Quantification in Latent Dynamics Models
- Videos are Sample-Efficient Supervisions: Behavior Cloning from Videos via Latent Representations
- InDRiVE: Reward-Free World-Model Pretraining for Autonomous Driving via Latent Disagreement
- STORM: Search-Guided Generative World Models for Robotic Manipulation
- Spectral Representation-based Reinforcement Learning
- Double Horizon Model-Based Policy Optimization
- World Models Can Leverage Human Videos for Dexterous Manipulation
- Latent Action World Models for Control with Unlabeled Trajectories
- Closing the Train-Test Gap in World Models for Gradient-Based Planning
- Prismatic World Model: Learning Compositional Dynamics for Planning in Hybrid Systems
- H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos
- Collaborative Causal Sensemaking: Closing the Complementarity Gap in Human-AI Decision Support
- On Memory: A comparison of memory mechanisms in world models
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds
- X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale
- World Models for Autonomous Navigation of Terrestrial Robots from LIDAR Observations
- Driving Beyond Privilege: Distilling Dense-Reward Knowledge into Sparse-Reward Policies
- Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
- Vehicle Dynamics Embedded World Models for Autonomous Driving
- From monoliths to modules: Decomposing transducers for efficient world modelling
- GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation
- SmallWorlds: Assessing Dynamics Understanding of World Models in Isolated Environments
- Embodied Intelligent Wireless (EIW): Synesthesia of Machines Empowered Wireless Communications
- From Discrete Plans to Real-World Execution: A World-Model-Driven Framework for Execution-Aware Multi-Agent Path Finding
- WPT: World-to-Policy Transfer via Online World Model Distillation
- Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?
- RynnVLA-002: A Unified Vision-Language-Action and World Model
- Counterfactual World Models via Digital Twin-conditioned Video Diffusion
- Latent Reasoning in TRMs is Secretly a Policy Improvement Operator
- IPR-1: Interactive Physical Reasoner
- Object-Centric World Models for Causality-Aware Reinforcement Learning
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning
- Measuring Progress in Deep Reinforcement Learning Sample Efficiency
- Reinforcement Learning through Active Inference
- Autonomous Vehicle Path Planning by Searching With Differentiable Simulation
- Simulating the Visual World with Artificial Intelligence: A Roadmap
- Reinforcement Learning Control of Quantum Error Correction
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- The LoCA Regret: A Consistent Metric to Evaluate Model-Based Behavior in Reinforcement Learning
- Representation Learning Enables Scalable Multitask Deep Reinforcement Learning
- Generative AI Meets 6G and Beyond: Diffusion Models for Semantic Communications
- Dynamic Sparsity: Challenging Common Sparsity Assumptions for Learning World Models in Robotic Reinforcement Learning Benchmarks
- ViPRA: Video Prediction for Robot Actions
- Learning to Focus: Prioritizing Informative Histories with Structured Attention Mechanisms in Partially Observable Reinforcement Learning
- When Object-Centric World Models Meet Policy Learning: From Pixels to Policies, and Where It Breaks
- DWM-RO: Decentralized World Models with Reasoning Offloading for SWIPT-enabled Satellite-Terrestrial HetNets
- Next-Latent Prediction Transformers Learn Compact World Models
- Scaling Agent Learning via Experience Synthesis
- Learning Interactive World Model for Object-Centric Reinforcement Learning
- Natural Building Blocks for Structured World Models: Theory, Evidence, and Scaling
- How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment
- Single-agent Reinforcement Learning Model for Regional Adaptive Traffic Signal Control
- Robust Single-Agent Reinforcement Learning for Regional Traffic Signal Control Under Demand Fluctuations
- Bootstrap Off-policy with World Model
- A Step Toward World Models: A Survey on Robotic Manipulation
- Learning Soft Robotic Dynamics with Active Exploration
- Constructing the Umwelt: Cognitive Planning through Belief-Intent Co-Evolution
- Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models
- Off-policy Reinforcement Learning with Model-based Exploration Augmentation
- Continual Model-Based Reinforcement Learning with Hypernetworks
- Action-Sufficient State Representation Learning for Control with Structural Constraints
- Centralized Model and Exploration Policy for Multi-Agent RL
- CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation
- Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method
- Model-Based Visual Planning with Self-Supervised Functional Distances
- Nautilus: From One Prompt to Plug-and-Play Robot Learning
- Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning
- Extracting Latent State Representations with Linear Dynamics from Rich Observations
- Zero-shot World Models Are Developmentally Efficient Learners
- World Simulation with Video Foundation Models for Physical AI
- TARC: Time-Adaptive Robotic Control
- Enhancing Tactile-based Reinforcement Learning for Robotic Control
- DreamerV3-XP: Optimizing exploration through uncertainty estimation
- Compositional Monte Carlo Tree Diffusion for Extendable Planning
- World Models Should Prioritize the Unification of Physical and Social Dynamics
- Real-Time Gait Adaptation for Quadrupeds using Model Predictive Control and Reinforcement Learning
- A Unified Framework for Zero-Shot Reinforcement Learning
- Evaluating Video Models as Simulators of Multi-Person Pedestrian Trajectories
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- LyTimeT: Towards Robust and Interpretable State-Variable Discovery
- Social World Model-Augmented Mechanism Design Policy Learning
- Semantic World Models
- Contrastive Variational Reinforcement Learning for Complex Observations
- Learning and Planning in Complex Action Spaces
- Efficient Model-Based Reinforcement Learning for Robot Control via Online Optimization
- A Comprehensive Survey on World Models for Embodied AI
- RM-RL: Role-Model Reinforcement Learning for Precise Robot Manipulation
- LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training
- SkyDreamer: Interpretable End-to-End Vision-Based Drone Racing with Model-Based Reinforcement Learning
- A Recipe for Efficient Sim-to-Real Transfer in Manipulation with Online Imitation-Pretrained World Models
- Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
- Context-Aware Model-Based Reinforcement Learning for Autonomous Racing
- FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasks
- MatterDoor: Sampling Zero-shot Spatio-semantic Priors using Generative Models
- R-WoM: Retrieval-augmented World Model For Computer-use Agents
- Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
- Ctrl-World: A Controllable Generative World Model for Robot Manipulation
- Analytical Survey of Learning with Low-Resource Data: From Analysis to Investigation
- Agent Learning via Early Experience
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
- UAMDP: Uncertainty-Aware Markov Decision Process for Risk-Constrained Reinforcement Learning from Probabilistic Forecasts
- Visual Perspective Taking for Opponent Behavior Modeling
- Dream to Recall: Imagination-Guided Experience Retrieval for Memory-Persistent Vision-and-Language Navigation
- Local Reinforcement Learning with Action-Conditioned Root Mean Squared Q-Functions
- What You Don't Know Can Hurt You: How Well do Latent Safety Filters Understand Partially Observable Safety Constraints?
- Learning to Crawl: Latent Model-Based Reinforcement Learning for Soft Robotic Adaptive Locomotion
- Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models
- Midway Network: Learning Representations for Recognition and Motion from Latent Dynamics
- DREAMer-VXS: A Latent World Model for Sample-Efficient AGV Exploration in Stochastic, Unobserved Environments
- Look-ahead Reasoning with a Learned Model in Imperfect Information Games
- Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees
- Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft
- TD-JEPA: Latent-predictive Representations for Zero-Shot Reinforcement Learning
- Can World Models Benefit VLMs for World Dynamics?
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- XQC: Well-conditioned Optimization Accelerates Deep Reinforcement Learning
- DyMoDreamer: World Modeling with Dynamic Modulation
- Training Agents Inside of Scalable World Models
- Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
- Structure in Deep Reinforcement Learning: A Survey and Open Problems
- VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
- WoW: Towards a World omniscient World model Through Embodied Interaction
- Context and Diversity Matter: The Emergence of In-Context Learning in World Models
- Latent Activation Editing: Inference-Time Refinement of Learned Policies for Safer Multirobot Navigation
- Embodied AI: From LLMs to World Models
- Robot Trajectron V2: A Probabilistic Shared Control Framework for Navigation
- DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions
- World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models
- Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
- TAPO: Transition-Aware Policy Optimization for LLM Agents
- AuricularWorld: Hierarchical Action-Guided World Modeling for Fine-Grained Auricular Structure Segmentation from CT Scans
- Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control
- QQWorld: Quantile-Quantile Matching for World Model Regularization
- MAD: Motion Appearance Decoupling for efficient Driving World Models
- Sampling Network Guided Cross-Entropy Method for Unsupervised Point Cloud Registration
- On the Importance of Multistability for Horizon Generalization in Reinforcement Learning
- Temporal Straightening for Latent Planning
- What Capable Agents Must Know: Selection Theorems for Robust Decision-Making under Uncertainty
- Latent Action Pretraining Through World Modeling
- Remote Sensing-Oriented World Model
- End2Race: Efficient End-to-End Imitation Learning for Real-Time F1Tenth Racing
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
- HERO: Hierarchical Extrapolation and Refresh for Efficient World Models
- SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models
- LaGarNet: Goal-Conditioned Recurrent State-Space Models for Pick-and-Place Garment Flattening
- TransZero: Parallel Tree Expansion in MuZero using Transformer Networks
- Action and Perception as Divergence Minimization
- Imagined Autocurricula
- Dream to Chat: Model-based Reinforcement Learning on Dialogues with User Belief Modeling
- Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
- PIANO: Physics Informed Autoregressive Network
- First Order Model-Based RL through Decoupled Backpropagation
- Convergence of regularized agent-state-based Q-learning in POMDPs
- Dynamics-Aligned Latent Imagination in Contextual World Models for Zero-Shot Generalization
- Learning from nature: insights into GraphDOP's representations of the Earth System
- WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
- Neural Robot Dynamics
- Categorical Policies: Multimodal Policy Learning and Exploration in Continuous Control
- Improving Pre-Trained Vision-Language-Action Policies with Model-Based Search
- MAPF-World: Action World Model for Multi-Agent Path Finding
- Visuomotor Grasping with World Models for Surgical Robots
- Learning State-Space Models of Dynamic Systems from Arbitrary Data using Joint Embedding Predictive Architectures
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- Integrating Reinforcement Learning with Visual Generative Models: Foundations and Advances
- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
- GoViG: Goal-Conditioned Visual Navigation Instruction Generation
- Reparameterization Proximal Policy Optimization
- Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-Distribution
- ME3-BEV: Mamba-Enhanced Deep Reinforcement Learning for End-to-End Autonomous Driving with BEV-Perception
- In-Context Reinforcement Learning via Communicative World Models
- DiWA: Diffusion Policy Adaptation with World Models
- VRPO: Rethinking Value Modeling for Robust RL Training under Noisy Supervision
- PIGDreamer: Privileged Information Guided World Models for Safe Partially Observable Reinforcement Learning
- Communicating Plans, Not Percepts: Scalable Multi-Agent Coordination with Embodied World Models
- Learning to Correspond Dynamical Systems
- Beyond Simulation: Benchmarking World Models for Planning and Causality in Autonomous Driving
Discussions
Related