Mastering Diverse Domains through World Models
2023/01/10 by Danijar Hafner, Hafner, Danijar, Jurgis Pasukonis +6 · 5 voices · 174 citations
Computer Science · #Anomaly Detection Techniques and Applications #Explainable Artificial Intelligence (XAI) #Machine Learning and Data Classification #cs.AI #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.2301.04104
openalex publication_date 2023/01/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
Abstract
Developing a general algorithm that learns to solve tasks across a wide range of applications has been a fundamental challenge in artificial intelligence. Although current reinforcement learning algorithms can be readily applied to tasks similar to what they have been developed for, configuring them for new application domains requires significant human expertise and experimentation. We present DreamerV3, a general algorithm that outperforms specialized methods across over 150 diverse tasks, with a single configuration. Dreamer learns a model of the environment and improves its behavior by imagining future scenarios. Robustness techniques based on normalization, balancing, and transformations enable stable learning across domains. Applied out of the box, Dreamer is the first algorithm to collect diamonds in Minecraft from scratch without human data or curricula. This achievement has been posed as a significant challenge in artificial intelligence that requires exploring farsighted strategies from pixels and sparse rewards in an open world. Our work allows solving challenging control problems without extensive experimentation, making reinforcement learning broadly applicable.
Cited by
- Robot-Factored World Models via Robot Rendering
- TRW: TRACE-RealWorld---An Auditable Consistency Contract for World Models as Materialized Views
- Persistent Computational State: A Session-Centric Runtime for Generative World Models
- Do World Action Models Generalize Better than VLAs? A Robustness Study
- The World Model Remembers, the Actor Forgets: Dream Rehearsal for Continual Model-Based RL
- Mobile Network Control with a World Model
- DWM: Separating World Effects from Actions in Latent World Models
- Multi-Horizon Consistency as Geometry: When Latent Dynamics Contract, and When They Do Not
- Predictive Training with Latent Imagination for Visual Quadruped Navigation
- World Translation: Minimizing Sim-to-Real Gap with Backward Dynamics Extraction and Unpaired Domain Translation
- SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning
- Planning as Emergent Behavior in Reinforcement Learning with Relational Hidden States
- Autoregressive Diffusion World Models for Off-Policy Evaluation of LLM Agents
- MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
- Adaptive Compute in Latent World Models: When Depth Helps, Hurts, or Doesn't Matter
- Orbis 2: A Hierarchical World Model for Driving
- Concept-Guided Spatial Regularization for World Models in Atari Pong
- Qwen-AgentWorld: Language World Models for General Agents
- Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners
- Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
- Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
- Dreamer-CDP: Improving Reconstruction-free World Models Via Continuous Deterministic Representation Prediction
- Cross-View World Models
- Learning Abstractions for Hierarchical Planning in Program-Synthesis Agents
- General agents contain world models
- CaRL: Learning Scalable Planning Policies with Simple Rewards
- Intrinsically-Motivated Humans and Agents in Open-World Exploration
- Synthesizing world models for bilevel planning
- Value-Based Deep RL Scales Predictably
- Stochastic Siamese MAE Pretraining for Longitudinal Medical Images
- Act2Goal: From World Model To General Goal-conditioned Policy
- LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments
- Action from Adjacent Set in Physical Space Outperforms the Best Prediction in World Models
- From World Models to World Action Models: A Concise Tutorial for Robotics
- DocHRL: A Hierarchical Reinforcement Learning Framework for Cost-Optimised Document Classification
- Biased Dreams: Limitations to Epistemic Uncertainty Quantification in Latent Dynamics Models
- Videos are Sample-Efficient Supervisions: Behavior Cloning from Videos via Latent Representations
- Recurrent Off-Policy Deep Reinforcement Learning Doesn't Have to be Slow
- InDRiVE: Reward-Free World-Model Pretraining for Autonomous Driving via Latent Disagreement
- From Word to World: Can Large Language Models be Implicit Text-based World Models?
- The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text
- Spectral Representation-based Reinforcement Learning
- Double Horizon Model-Based Policy Optimization
- DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos
- Scaling Behavior of Discrete Diffusion Language Models
- Closing the Train-Test Gap in World Models for Gradient-Based Planning
- Optimal Perturbation Budget Allocation for Data Poisoning in Offline Reinforcement Learning
- F2: Offline Reinforcement Learning for Hamiltonian Simulation via Free-Fermionic Subroutine Compilation
- Learning Without Time-Based Embodiment Resets in Soft-Actor Critic
- Generative Recursive Reasoning
- Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism
- World Models for Autonomous Navigation of Terrestrial Robots from LIDAR Observations
- Driving Beyond Privilege: Distilling Dense-Reward Knowledge into Sparse-Reward Policies
- Vehicle Dynamics Embedded World Models for Autonomous Driving
- Objects in Generated Videos Are Slower Than They Appear: Models Suffer Sub-Earth Gravity and Don't Know Galileo's Principle...for now
- GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation
- From Regression to Classification: Exploring the Benefits of Categorical Representations of Energy in MLIPs
- Goal-Driven Reward by Video Diffusion Models for Reinforcement Learning
- Language-conditioned world model improves policy generalization by reading environmental descriptions
- Subgoal Graph-Augmented Planning for LLM-Guided Open-World Reinforcement Learning
- SOMBRL: Scalable and Optimistic Model-Based RL
- WPT: World-to-Policy Transfer via Online World Model Distillation
- Learning Massively Multitask World Models for Continuous Control
- Weakly-supervised Latent Models for Task-specific Visual-Language Control
- Dreaming Falcon: Physics-Informed Model-Based Reinforcement Learning for Quadcopters
- Counterfactual World Models via Digital Twin-conditioned Video Diffusion
- IPR-1: Interactive Physical Reasoner
- ARCTraj: A Dataset and Benchmark of Human Reasoning Trajectories for Abstract Problem Solving
- Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
- Simulating the Visual World with Artificial Intelligence: A Roadmap
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- PADiff: Predictive and Adaptive Diffusion Policies for Ad Hoc Teamwork
- MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios
- DWM-RO: Decentralized World Models with Reasoning Offloading for SWIPT-enabled Satellite-Terrestrial HetNets
- WorldPlanner: Monte Carlo Tree Search and MPC with Action-Conditioned Visual World Models
- Learning Interactive World Model for Object-Centric Reinforcement Learning
- Single-agent Reinforcement Learning Model for Regional Adaptive Traffic Signal Control
- Robust Single-Agent Reinforcement Learning for Regional Traffic Signal Control Under Demand Fluctuations
- Bootstrap Off-policy with World Model
- Learning Generalizable Visuomotor Policy through Dynamics-Alignment
- A Step Toward World Models: A Survey on Robotic Manipulation
- Off-policy Reinforcement Learning with Model-based Exploration Augmentation
- ActSWM: Action-Sensitive World Models for Long-Horizon Planning in Open-World Games
- Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels
- D2 Actor Critic: Diffusion Actor Meets Distributional Critic
- Nautilus: From One Prompt to Plug-and-Play Robot Learning
- Large Emotional World Model
- Human Machine Social Hybrid Intelligence:A Collaborative Decision Making Framework for Large Model Agent Groups and Human Experts
- A Survey on Efficient Vision-Language-Action Models
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
- Semantic Communications with World Models
- Guardian: Decoupling Exploration from Safety in Reinforcement Learning
- I2-NeRF: Learning Neural Radiance Fields Under Physically-Grounded Media Interactions
- DreamerV3-XP: Optimizing exploration through uncertainty estimation
- World Models Should Prioritize the Unification of Physical and Social Dynamics
- X-Ego: Acquiring Team-Level Tactical Situational Awareness via Cross-Egocentric Contrastive Video Representation Learning
- Social World Model-Augmented Mechanism Design Policy Learning
- Socialized Learning and Emergent Behaviors in Multi-Agent Systems based on Multimodal Large Language Models
- Efficient Model-Based Reinforcement Learning for Robot Control via Online Optimization
- Heterogeneous Adversarial Play in Interactive Environments
- Learning to play: A Multimodal Agent for 3D Game-Play
- A Recipe for Efficient Sim-to-Real Transfer in Manipulation with Online Imitation-Pretrained World Models
- Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents
- Fixing That Free Lunch: When, Where, and Why Synthetic Data Fails in Model-Based Policy Optimization
- CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving
- Rethinking the Role of Dynamic Sparse Training for Scalable Deep Reinforcement Learning
- One Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided Exploration
- Ego-Vision World Model for Humanoid Contact Planning
- Context-Aware Model-Based Reinforcement Learning for Autonomous Racing
- FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasks
- R-WoM: Retrieval-augmented World Model For Computer-use Agents
- Expressive Value Learning for Scalable Offline Reinforcement Learning
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
- Learning to Crawl: Latent Model-Based Reinforcement Learning for Soft Robotic Adaptive Locomotion
- BuilderBench -- A benchmark for generalist agents
- Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models
- MultiModal Action Conditioned Video Generation
- SITCOM: Scaling Inference-Time COMpute for VLAs
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- Noise-Guided Transport for Imitation Learning
- World Model for AI Autonomous Navigation in Mechanical Thrombectomy
- DyMoDreamer: World Modeling with Dynamic Modulation
- Unifying Agent Interaction and World Information for Multi-agent Coordination
- Dual-Scale World Models for LLM Agents Towards Hard-Exploration Problems
- Adversarial Diffusion for Robust Reinforcement Learning
- WoW: Towards a World omniscient World model Through Embodied Interaction
- FlowDrive: moderated flow matching with data balancing for trajectory planning
- Model-Based Reinforcement Learning under Random Observation Delays
- Wonder Wins Ways: Curiosity-Driven Exploration through Multi-Agent Contextual Calibration
- Embodied AI: From LLMs to World Models
- ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
- World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models
- Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
- TAPO: Transition-Aware Policy Optimization for LLM Agents
- ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow
- What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations
- Temporal Straightening for Latent Planning
- What Capable Agents Must Know: Selection Theorems for Robust Decision-Making under Uncertainty
- Latent Action Pretraining Through World Modeling
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
- HERO: Hierarchical Extrapolation and Refresh for Efficient World Models
- Foundation Models as World Models: A Foundational Study in Text-Based GridWorlds
- Designing Latent Safety Filters using Pre-Trained Vision Models
- Pre-trained Visual Representations Generalize Where it Matters in Model-Based Reinforcement Learning
- OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
- LaGarNet: Goal-Conditioned Recurrent State-Space Models for Pick-and-Place Garment Flattening
- How well can LLMs provide planning feedback in grounded environments?
- A Data-Driven Discretized CS:GO Simulation Environment to Facilitate Strategic Multi-Agent Planning Research
- Offline vs. Online Learning in Model-based RL: Lessons for Data Collection Strategies
- COGITAO: A Visual Reasoning Framework To Study Compositionality & Generalization
- Planning with Reasoning using Vision Language World Model
- STORI: A Benchmark and Taxonomy for Stochastic Environments
- First Order Model-Based RL through Decoupled Backpropagation
- NiceWebRL: a Python library for human subject experiments with reinforcement learning environments
- WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
- Neural Robot Dynamics
- Categorical Policies: Multimodal Policy Learning and Exploration in Continuous Control
- CrafterDojo: A Suite of Foundation Models for Building Open-Ended Embodied Agents in Crafter
- Contrastive Representations for Temporal Reasoning
- SIGN: Safety-Aware Image-Goal Navigation for Autonomous Drones via Reinforcement Learning
- Finite Automata Extraction: Low-data World Model Learning as Programs from Gameplay Video
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- Integrating Reinforcement Learning with Visual Generative Models: Foundations and Advances
- GoViG: Goal-Conditioned Visual Navigation Instruction Generation
- Playing Atari Space Invaders with Sparse Cosine Optimized Policy Evolution
- Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction
- DiWA: Diffusion Policy Adaptation with World Models
- PIGDreamer: Privileged Information Guided World Models for Safe Partially Observable Reinforcement Learning
- Communicating Plans, Not Percepts: Scalable Multi-Agent Coordination with Embodied World Models
- GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration
- One-Step Flow Policy Mirror Descent
- Machine learning in video games [wikipedia]
- January–March 2023 in science [wikipedia]
- Timeline of computing 2020–present [wikipedia]
Discussions
- Mastering diverse domains through world models [hn, 76 points, 7 comments]
- DeepMind's new AI ("DreamerV3") finds diamonds in Minecraft without being taught [lemmy, 22 points, 10 comments]
- I've seen people struggle to get good performance out of latent replay. But that being said, there has been some success in latent replay more recently, notably, the Dreamer models from Danijar Hafner [bsky, 2 points, 1 comments]
- All of these works frame world model as a bayesian inference of the hidden markovian state from observations and actions (using sequential variational inference). They build on Dreamer (arxiv.org/abs [bsky, 1 points, 0 comments]
- LLMs can't do this, but the Dreamer class of networks absolutely can. arxiv.org/abs/2301.04104 [bsky, 0 points, 0 comments]
Related