DayDreamer: World Models for Physical Robot Learning
2022/06/28 by Philipp Wu, Alejandro Escontrela, Wu, Philipp +7 · 1 voice · 114 citations
Computer Science · Engineering · #Artificial intelligence #Computer science #Computer vision #Engineering #Human Pose and Action Recognition #Human–computer interaction #Limiting #Machine Learning and Data Classification #Mobile robot #Reinforcement Learning in Robotics #Reinforcement learning #Robot #Robot learning #Robotics #Simulation #Software deployment #cs.AI #cs.LG #cs.RO
paper · pdf · doi:10.48550/arxiv.2206.14176
published in arXiv (Cornell University) (Cornell University) · Website: https://danijar.com/daydreamer
arxiv created 2022/06/28 · openalex publication_date 2022/06/28 · arxiv published 2022/06/28 · arxiv updated 2022/06/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
To solve tasks in complex environments, robots need to learn from experience. Deep reinforcement learning is a common approach to robot learning but requires a large amount of trial and error to learn, limiting its deployment in the physical world. As a consequence, many advances in robot learning rely on simulators. On the other hand, learning inside of simulators fails to capture the complexity of the real world, is prone to simulator inaccuracies, and the resulting behaviors do not adapt to changes in the world. The Dreamer algorithm has recently shown great promise for learning from small amounts of interaction by planning within a learned world model, outperforming pure reinforcement learning in video games. Learning a world model to predict the outcomes of potential actions enables planning in imagination, reducing the amount of trial and error needed in the real environment. However, it is unknown whether Dreamer can facilitate faster learning on physical robots. In this paper, we apply Dreamer to 4 robots to learn online and directly in the real world, without simulators. Dreamer trains a quadruped robot to roll off its back, stand up, and walk from scratch and without resets in only 1 hour. We then push the robot and find that Dreamer adapts within 10 minutes to withstand perturbations or quickly roll over and stand back up. On two different robotic arms, Dreamer learns to pick and place multiple objects directly from camera images and sparse rewards, approaching human performance. On a wheeled robot, Dreamer learns to navigate to a goal position purely from camera images, automatically resolving ambiguity about the robot orientation. Using the same hyperparameters across all experiments, we find that Dreamer is capable of online learning in the real world, establishing a strong baseline. We release our infrastructure for future applications of world models to robot learning.
Cited by
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- WCM: World-Cognition Model for Generalizable Human-Robot Interaction
- From World Models to World Action Models: A Concise Tutorial for Robotics
- PhysFire-WM: A Physics-Informed World Model for Emulating Fire Spread Dynamics
- Self-motion as a structural prior for coherent and robust formation of cognitive maps
- ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning
- DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos
- ARCADE: Adaptive Robot Control with Online Changepoint-Aware Bayesian Dynamics Learning
- World Models for Autonomous Navigation of Terrestrial Robots from LIDAR Observations
- GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation
- H-Zero: Cross-Humanoid Locomotion Pretraining Enables Few-shot Novel Embodiment Transfer
- SOMBRL: Scalable and Optimistic Model-Based RL
- WPT: World-to-Policy Transfer via Online World Model Distillation
- Weakly-supervised Latent Models for Task-specific Visual-Language Control
- Dreaming Falcon: Physics-Informed Model-Based Reinforcement Learning for Quadcopters
- Counterfactual World Models via Digital Twin-conditioned Video Diffusion
- ViPRA: Video Prediction for Robot Actions
- Learning Interactive World Model for Object-Centric Reinforcement Learning
- A Step Toward World Models: A Survey on Robotic Manipulation
- Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method
- Nautilus: From One Prompt to Plug-and-Play Robot Learning
- Zero-shot World Models Are Developmentally Efficient Learners
- Deep Active Inference with Diffusion Policy and Multiple Timescale World Model for Real-World Exploration and Navigation
- I2-NeRF: Learning Neural Radiance Fields Under Physically-Grounded Media Interactions
- World Models Should Prioritize the Unification of Physical and Social Dynamics
- Efficient Model-Based Reinforcement Learning for Robot Control via Online Optimization
- A Comprehensive Survey on World Models for Embodied AI
- SkyDreamer: Interpretable End-to-End Vision-Based Drone Racing with Model-Based Reinforcement Learning
- CADRE: Dynamic Catching via Implicit Contact Descriptors and Task-Appropriate Recovery Affordances
- Constraint-Aware Reinforcement Learning via Adaptive Action Scaling
- Ctrl-World: A Controllable Generative World Model for Robot Manipulation
- Learning to Crawl: Latent Model-Based Reinforcement Learning for Soft Robotic Adaptive Locomotion
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- Act to See, See to Act: Diffusion-Driven Perception-Action Interplay for Adaptive Policies
- Training Agents Inside of Scalable World Models
- Context and Diversity Matter: The Emergence of In-Context Learning in World Models
- CORE: Full-Path Evaluation of LLM Agents Beyond Final State
- Embodied AI: From LLMs to World Models
- Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
- Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control
- Real-Time Reinforcement Learning for Dynamic Tasks with a Parallel Soft Robot
- Remote Sensing-Oriented World Model
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
- HERO: Hierarchical Extrapolation and Refresh for Efficient World Models
- Designing Latent Safety Filters using Pre-Trained Vision Models
- SHaRe-RL: Structured, Interactive Reinforcement Learning for Contact-Rich Industrial Assembly Tasks
- Empowering Multi-Robot Cooperation via Sequential World Models
- InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation
- Leveraging AI Agents for Autonomous Networks: A Reference Architecture and Empirical Studies
- Language Self-Play For Data-Free Training
- Towards bridging the gap: Systematic sim-to-real transfer for diverse legged robots
- PlayerOne: Egocentric World Simulator
- First Order Model-Based RL through Decoupled Backpropagation
- From Tabula Rasa to Emergent Abilities: Discovering Robot Skills via Real-World Unsupervised Quality-Diversity
- Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids
- Human Centric General Physical Intelligence for Agile Manufacturing Automation
- Multi-Group Equivariant Augmentation for Reinforcement Learning in Robot Manipulation
- Visuomotor Grasping with World Models for Surgical Robots
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- How Safe Will I Be Given What I Saw? Calibrated Prediction of Safety Chances for Image-Controlled Autonomy
- MonoMPC: Monocular Vision Based Navigation with Learned Collision Model and Risk-Aware Model Predictive Control
- Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-Distribution
- Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
- Achieving Precise and Reliable Locomotion with Differentiable Simulation-Based System Identification
- ActionSink: Toward Precise Robot Manipulation with Dynamic Integration of Action Flow
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Time-Aware World Model for Adaptive Prediction and Control
- NeuralOS: Towards Simulating Operating Systems via Neural Generative Models
- AirScape: An Aerial Generative World Model with Motion Controllability
- Spatial-Temporal Aware Visuomotor Diffusion Policy Learning
- Advances, challenges, and opportunities for legged robots
- A Survey: Learning Embodied Intelligence from Physical Simulators and World Models
- Memory Allocation in Resource-Constrained Reinforcement Learning
- NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments
- Goal-VLA: Image-Generative VLMs as Object-Centric World Models Empowering Zero-shot Robot Manipulation
- Unified Vision-Language-Action Model
- ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model
- Efficient Generation of Diverse Cooperative Agents with World Models
- MinD: Learning A Dual-System World Model for Real-Time Planning and Implicit Risk Analysis
- Reimagination with Test-time Observation Interventions: Distractor-Robust World Model Predictions for Visual Model Predictive Control
- From Pixels to CSI: Distilling Latent Dynamics For Efficient Wireless Resource Management
- A Survey on Imitation Learning for Contact-Rich Tasks in Robotics
- DynaGuide: Steering Diffusion Polices with Active Dynamic Guidance
- A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to Search
- Video World Models with Long-term Spatial Memory
- WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
- Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning
- Variational Adaptive Noise and Dropout towards Stable Recurrent Neural Networks
- WoMAP: World Models For Embodied Open-Vocabulary Object Localization
- Learning agile soccer skills for a bipedal robot with deep reinforcement learning
- StateSpaceDiffuser: Bringing Long Context to Diffusion World Models
- Long-Context State-Space Video World Models
- OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory Generation
- DriveX: Omni Scene Modeling for Learning Generalizable World Knowledge in Autonomous Driving
- WorldEval: World Model as Real-World Robot Policies Evaluator
- DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion
- Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
- Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
- FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation
- Drive Fast, Learn Faster: On-Board RL for High Performance Autonomous Racing
- LineFlow: A Framework to Learn Active Control of Production Lines
- Learning Local Causal World Models with State Space Models and Attention
- Enhancing Policy Learning with World-Action Model
- Quo Vadis, World Modeling?
- World model-driven process industry operations: An offline reinforcement learning solution based on conditional diffusion
- Agile legged locomotion in reconfigurable modular robots
- VLANeXt: Recipes for Building Strong VLA Models
- Causal World Modeling for Robot Control
- WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation
- PIN-WM: Learning Physics-INformed World Models for Non-Prehensile Manipulation
- Latent Diffusion Planning for Imitation Learning
- Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination
- Sim-to-Real of Humanoid Locomotion Policies via Joint Torque Space Perturbation Injection
- Agent-Arena: A General Framework for Evaluating Control Algorithms
Discussions
Related