Dream to Control: Learning Behaviors by Latent Imagination
2019/12/03 by Danijar Hafner, Hafner, Danijar, Timothy Lillicrap +5 · 2 voices · 366 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics #cs.AI #cs.LG #cs.RO
paper · pdf · doi:10.48550/arxiv.1912.01603
9 pages, 12 figures
arxiv published 2019/12/03 · arxiv created 2020/03/17 · arxiv updated 2020/03/18
Abstract
Learned world models summarize an agent's experience to facilitate learning complex behaviors. While learning world models from high-dimensional sensory inputs is becoming feasible through deep learning, there are many potential ways for deriving behaviors from them. We present Dreamer, a reinforcement learning agent that solves long-horizon tasks from images purely by latent imagination. We efficiently learn behaviors by propagating analytic gradients of learned state values back through trajectories imagined in the compact state space of a learned world model. On 20 challenging visual control tasks, Dreamer exceeds existing approaches in data-efficiency, computation time, and final performance.
Citations
Cited by
- Act2Goal: From World Model To General Goal-conditioned Policy
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback
- Action from Adjacent Set in Physical Space Outperforms the Best Prediction in World Models
- N0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation
- FeelWorld: Visuo-Tactile World Model for Hierarchical Contact Prediction and Planning
- Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation
- Medical world models in healthcare: foundations, applications, and challenges for trustworthy clinical translation
- False Prophets: On the Security of World Models in Agentic Systems
- PatchWorld: Gradient-Free Optimization of Executable World Models for Agent Environments
- Bridging MARL to SARL: An Order-Independent Multi-Agent Transformer via Latent Consensus
- Biased Dreams: Limitations to Epistemic Uncertainty Quantification in Latent Dynamics Models
- Videos are Sample-Efficient Supervisions: Behavior Cloning from Videos via Latent Representations
- InDRiVE: Reward-Free World-Model Pretraining for Autonomous Driving via Latent Disagreement
- STORM: Search-Guided Generative World Models for Robotic Manipulation
- Spectral Representation-based Reinforcement Learning
- Double Horizon Model-Based Policy Optimization
- World Models for Learning Dexterous Hand-Object Interactions from Human Videos
- Latent Action World Models for Control with Unlabeled Trajectories
- Closing the Train-Test Gap in World Models for Gradient-Based Planning
- Prismatic World Model: Learning Compositional Dynamics for Planning in Hybrid Systems
- H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos
- Collaborative Causal Sensemaking: Closing the Complementarity Gap in Human-AI Decision Support
- On Memory: A comparison of memory mechanisms in world models
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds
- X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale
- World Models for Autonomous Navigation of Terrestrial Robots from LIDAR Observations
- Driving Beyond Privilege: Distilling Dense-Reward Knowledge into Sparse-Reward Policies
- Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
- Vehicle Dynamics Embedded World Models for Autonomous Driving
- From monoliths to modules: Decomposing transducers for efficient world modelling
- GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation
- SmallWorlds: Assessing Dynamics Understanding of World Models in Isolated Environments
- Embodied Intelligent Wireless (EIW): Synesthesia of Machines Empowered Wireless Communications
- From Discrete Plans to Real-World Execution: A World-Model-Driven Framework for Execution-Aware Multi-Agent Path Finding
- WPT: World-to-Policy Transfer via Online World Model Distillation
- Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?
- RynnVLA-002: A Unified Vision-Language-Action and World Model
- Counterfactual World Models via Digital Twin-conditioned Video Diffusion
- Latent Reasoning in TRMs is Secretly a Policy Improvement Operator
- IPR-1: Interactive Physical Reasoner
- Object-Centric World Models for Causality-Aware Reinforcement Learning
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning
- Measuring Progress in Deep Reinforcement Learning Sample Efficiency
- Reinforcement Learning through Active Inference
- Autonomous Vehicle Path Planning by Searching With Differentiable Simulation
- Simulating the Visual World with Artificial Intelligence: A Roadmap
- Reinforcement Learning Control of Quantum Error Correction
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- The LoCA Regret: A Consistent Metric to Evaluate Model-Based Behavior in Reinforcement Learning
- Representation Learning Enables Scalable Multitask Deep Reinforcement Learning
- Generative AI Meets 6G and Beyond: Diffusion Models for Semantic Communications
- Dynamic Sparsity: Challenging Common Sparsity Assumptions for Learning World Models in Robotic Reinforcement Learning Benchmarks
- ViPRA: Video Prediction for Robot Actions
- Learning to Focus: Prioritizing Informative Histories with Structured Attention Mechanisms in Partially Observable Reinforcement Learning
- When Object-Centric World Models Meet Policy Learning: From Pixels to Policies, and Where It Breaks
- DWM-RO: Decentralized World Models with Reasoning Offloading for SWIPT-enabled Satellite-Terrestrial HetNets
- Next-Latent Prediction Transformers Learn Compact World Models
- Scaling Agent Learning via Experience Synthesis
- Learning Interactive World Model for Object-Centric Reinforcement Learning
- Natural Building Blocks for Structured World Models: Theory, Evidence, and Scaling
- How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment
- Single-agent Reinforcement Learning Model for Regional Adaptive Traffic Signal Control
- Robust Single-Agent Reinforcement Learning for Regional Traffic Signal Control Under Demand Fluctuations
- Bootstrap Off-policy with World Model
- A Step Toward World Models: A Survey on Robotic Manipulation
- Learning Soft Robotic Dynamics with Active Exploration
- Constructing the Umwelt: Cognitive Planning through Belief-Intent Co-Evolution
- Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models
- Off-policy Reinforcement Learning with Model-based Exploration Augmentation
- Continual Model-Based Reinforcement Learning with Hypernetworks
- Action-Sufficient State Representation Learning for Control with Structural Constraints
- Centralized Model and Exploration Policy for Multi-Agent RL
- CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation
- Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method
- Model-Based Visual Planning with Self-Supervised Functional Distances
- Nautilus: From One Prompt to Plug-and-Play Robot Learning
- Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning
- Extracting Latent State Representations with Linear Dynamics from Rich Observations
- Zero-shot World Models Are Developmentally Efficient Learners
- World Simulation with Video Foundation Models for Physical AI
- TARC: Time-Adaptive Robotic Control
- Enhancing Tactile-based Reinforcement Learning for Robotic Control
- DreamerV3-XP: Optimizing exploration through uncertainty estimation
- Compositional Monte Carlo Tree Diffusion for Extendable Planning
- World Models Should Prioritize the Unification of Physical and Social Dynamics
- Real-Time Gait Adaptation for Quadrupeds using Model Predictive Control and Reinforcement Learning
- A Unified Framework for Zero-Shot Reinforcement Learning
- PEDRA: Evaluating the Realism of Pedestrian Dynamics in Video Generation
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- LyTimeT: Towards Robust and Interpretable State-Variable Discovery
- Social World Model-Augmented Mechanism Design Policy Learning
- Semantic World Models
- Contrastive Variational Reinforcement Learning for Complex Observations
- Learning and Planning in Complex Action Spaces
- Efficient Model-Based Reinforcement Learning for Robot Control via Online Optimization
- A Comprehensive Survey on World Models for Embodied AI
- RM-RL: Role-Model Reinforcement Learning for Precise Robot Manipulation
- LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training
- SkyDreamer: Interpretable End-to-End Vision-Based Drone Racing with Model-Based Reinforcement Learning
- A Recipe for Efficient Sim-to-Real Transfer in Manipulation with Online Imitation-Pretrained World Models
- Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
- Context-Aware Model-Based Reinforcement Learning for Autonomous Racing
- FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasks
- MatterDoor: Sampling Zero-shot Spatio-semantic Priors using Generative Models
- R-WoM: Retrieval-augmented World Model For Computer-use Agents
- Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
- Ctrl-World: A Controllable Generative World Model for Robot Manipulation
- Analytical Survey of Learning with Low-Resource Data: From Analysis to Investigation
- Agent Learning via Early Experience
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
- UAMDP: Uncertainty-Aware Markov Decision Process for Risk-Constrained Reinforcement Learning from Probabilistic Forecasts
- Visual Perspective Taking for Opponent Behavior Modeling
- Dream to Recall: Imagination-Guided Experience Retrieval for Memory-Persistent Vision-and-Language Navigation
- Local Reinforcement Learning with Action-Conditioned Root Mean Squared Q-Functions
- How Well Do Latent World Models Understand Partially Observable Safety Constraints?
- Learning to Crawl: Latent Model-Based Reinforcement Learning for Soft Robotic Adaptive Locomotion
- Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models
- Midway Network: Learning Representations for Recognition and Motion from Latent Dynamics
- DREAMer-VXS: A Latent World Model for Sample-Efficient AGV Exploration in Stochastic, Unobserved Environments
- Look-ahead Reasoning with a Learned Model in Imperfect Information Games
- Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees
- Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft
- TD-JEPA: Latent-predictive Representations for Zero-Shot Reinforcement Learning
- Can World Models Benefit VLMs for World Dynamics?
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- XQC: Well-conditioned Optimization Accelerates Deep Reinforcement Learning
- DyMoDreamer: World Modeling with Dynamic Modulation
- Training Agents Inside of Scalable World Models
- Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
- Structure in Deep Reinforcement Learning: A Survey and Open Problems
- VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
- WoW: Towards a World omniscient World model Through Embodied Interaction
- Context and Diversity Matter: The Emergence of In-Context Learning in World Models
- Latent Activation Editing: Inference-Time Refinement of Learned Policies for Safer Multirobot Navigation
- Embodied AI: From LLMs to World Models
- Robot Trajectron V2: A Probabilistic Shared Control Framework for Navigation
- DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions
- World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models
- Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
- TAPO: Transition-Aware Policy Optimization for LLM Agents
- AuricularWorld: Hierarchical Action-Guided World Modeling for Fine-Grained Auricular Structure Segmentation from CT Scans
- Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control
- QQWorld: Quantile-Quantile Matching for World Model Regularization
- MAD: Motion Appearance Decoupling for efficient Driving World Models
- Sampling Network Guided Cross-Entropy Method for Unsupervised Point Cloud Registration
- On the Importance of Multistability for Horizon Generalization in Reinforcement Learning
- Temporal Straightening for Latent Planning
- What Capable Agents Must Know: Selection Theorems for Robust Decision-Making under Uncertainty
- Latent Action Pretraining Through World Modeling
- Remote Sensing-Oriented World Model
- End2Race: Efficient End-to-End Imitation Learning for Real-Time F1Tenth Racing
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
- HERO: Hierarchical Extrapolation and Refresh for Efficient World Models
- SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models
- LaGarNet: Goal-Conditioned Recurrent State-Space Models for Pick-and-Place Garment Flattening
- TransZero: Parallel Tree Expansion in MuZero using Transformer Networks
- Action and Perception as Divergence Minimization
- Imagined Autocurricula
- Dream to Chat: Model-based Reinforcement Learning on Dialogues with User Belief Modeling
- Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
- PIANO: Physics Informed Autoregressive Network
- PlayerOne: Egocentric World Simulator
- First Order Model-Based RL through Decoupled Backpropagation
- Convergence of regularized agent-state-based Q-learning in POMDPs
- Dynamics-Aligned Latent Imagination in Contextual World Models for Zero-Shot Generalization
- Learning from nature: insights into GraphDOP's representations of the Earth System
- WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
- Neural Robot Dynamics
- Categorical Policies: Multimodal Policy Learning and Exploration in Continuous Control
- Improving Pre-Trained Vision-Language-Action Policies with Model-Based Search
- MAPF-World: Action World Model for Multi-Agent Path Finding
- Visuomotor Grasping with World Models for Surgical Robots
- Learning State-Space Models of Dynamic Systems from Arbitrary Data using Joint Embedding Predictive Architectures
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- Integrating Reinforcement Learning with Visual Generative Models: Foundations and Advances
- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
- GoViG: Goal-Conditioned Visual Navigation Instruction Generation
- Reparameterization Proximal Policy Optimization
- Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-Distribution
- ME3-BEV: Mamba-Enhanced Deep Reinforcement Learning for End-to-End Autonomous Driving with BEV-Perception
- In-Context Reinforcement Learning via Communicative World Models
- ReSim: Reliable World Simulation for Autonomous Driving
- DiWA: Diffusion Policy Adaptation with World Models
- VRPO: Rethinking Value Modeling for Robust RL Training under Noisy Supervision
- PIGDreamer: Privileged Information Guided World Models for Safe Partially Observable Reinforcement Learning
- Communicating Plans, Not Percepts: Scalable Multi-Agent Coordination with Embodied World Models
- Learning to Correspond Dynamical Systems
- Beyond Simulation: Benchmarking World Models for Planning and Causality in Autonomous Driving
- Back to the Features: DINO as a Foundation for Video World Models
- Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models
- Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories
- Neuro-Inspired Inverse Learning for Planning and Control
- High Performance Across Two Atari Paddle Games Using the Same Perceptual Control Architecture Without Training
- Variational State-Space Models for Localisation and Dense 3D Mapping in 6 DoF
- Reinforce Lifelong Interaction Value of User-Author Pairs for Large-Scale Recommendation Systems
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Deep Reinforcement and InfoMax Learning
- Context-Aware Safe Reinforcement Learning for Non-Stationary Environments
- Intervention Design for Effective Sim2Real Transfer
- Time-Aware World Model for Adaptive Prediction and Control
- Diffusion-Based Imaginative Coordination for Bimanual Manipulation
- Taming generative video models for zero-shot optical flow extraction
- DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search
- AquaJEPA: Action-Conditioned Multimodal Predictive Representations for Underwater Robot Dynamics
- Spatial-Temporal Aware Visuomotor Diffusion Policy Learning
- Sample-Efficient Reinforcement Learning Controller for Deep Brain Stimulation in Parkinson's Disease
- FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution
- DARIL: When Imitation Learning outperforms Reinforcement Learning in Surgical Action Planning
- Accurate and Efficient World Modeling with Masked Latent Transformers
- NRSeg: Noise-Resilient Learning for BEV Semantic Segmentation via Driving World Models
- Dyn-O: Building Structured World Models with Object-Centric Representations
- Mitigating Goal Misgeneralization via Minimax Regret
- ActionParty: Multi-Subject Action Binding in Generative Video Games
- TD-MPC-Opt: Distilling Model-Based Multi-Task Reinforcement Learning Agents
- A Survey: Learning Embodied Intelligence from Physical Simulators and World Models
- NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments
- Goal-VLA: Image-Generative VLMs as Object-Centric World Models Empowering Zero-shot Robot Manipulation
- Curious Causality-Seeking Agents Learn Meta Causal World
- ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation
- Simple and Effective VAE Training with Calibrated Decoders
- SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model
- Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
- Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation
- rQdia: Regularizing Q-Value Distributions With Image Augmentation
- Whole-Body Conditioned Egocentric Video Prediction
- From 2D to 3D Cognition: A Brief Survey of General World Models
- Learning Task Belief Similarity with Latent Dynamics for Meta-Reinforcement Learning
- Is an object-centric representation beneficial for robotic manipulation ?
- FlightKooba: A Fast Interpretable FTP Model
- Unified Vision-Language-Action Model
- ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model
- Efficient Generation of Diverse Cooperative Agents with World Models
- Transformer World Model for Sample Efficient Multi-Agent Reinforcement Learning
- Adapting Vision-Language Models for Evaluating World Models
- TransDreamerV3: Implanting Transformer In DreamerV3
- Measuring Intent Comprehension in LLMs
- Zero-Shot Reinforcement Learning Under Partial Observability
- GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation
- A Survey on World Models Grounded in Acoustic Physical Information
- DynaGuide: Steering Diffusion Polices with Active Dynamic Guidance
- Revealing the Challenges of Sim-to-Real Transfer in Model-Based Reinforcement Learning via Latent Space Modeling
- TARDIS STRIDE: A Spatio-Temporal Road Image Dataset and World Model for Autonomy
- Fast Monte Carlo Tree Diffusion: 100x Speedup via Parallel Sparse Planning
- WorldLLM: Improving LLMs' world modeling using curiosity-driven theory-making
- Action-Adaptive Continual Learning: Enabling Policy Generalization under Dynamic Action Spaces
- A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to Search
- Model-Based Offline Planning
- Video World Models with Long-term Spatial Memory
- Dream to Generalize: Zero-Shot Model-Based Reinforcement Learning for Unseen Visual Distractions
- WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
- seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models
- Simple, Good, Fast: Self-Supervised World Models Free of Baggage
- Data-assimilated model-informed reinforcement learning
- NoiseAR: AutoRegressing Initial Noise Prior for Diffusion Models
- Sparse Imagination for Efficient Visual World Model Planning
- WoMAP: World Models For Embodied Open-Vocabulary Object Localization
- LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model
- MAGIK: Mapping to Analogous Goals via Imagination-enabled Knowledge Transfer
- WorldGym: World Model as An Environment for Policy Evaluation
- AXIOM: Learning to Play Games in Minutes with Expanding Object-Centric Models
- R3DM: Enabling Role Discovery and Diversity Through Dynamics Models in Multi-agent Reinforcement Learning
- LoopNav: Benchmarking Spatial Consistency in World Models
- StateSpaceDiffuser: Bringing Long Context to Diffusion World Models
- Calibrated Value-Aware Model Learning with Probabilistic Environment Models
- Revisiting Multi-Agent World Modeling from a Diffusion-Inspired Perspective
- Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
- Long-Context State-Space Video World Models
- medDreamer: Model-Based Reinforcement Learning with Latent Imagination on Complex EHRs for Clinical Decision Support
- Situationally-Aware Dynamics Learning
- DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
- OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory Generation
- Deep Active Inference Agents for Delayed and Long-Horizon Environments
- DriveX: Omni Scene Modeling for Learning Generalizable World Knowledge in Autonomous Driving
- ProWorld: Progress-Aware Hyperbolic World Models for Long-Horizon Visual Goal Reaching
- Bayesian Meta-Reinforcement Learning with Laplace Variational Recurrent Networks
- ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos
- Is Single-View Mesh Reconstruction Ready for Robotics?
- FedWorld: Scope-Aware Federation of Agent World Models
- Imagine Beyond! Distributionally Robust Auto-Encoding for State Space Coverage in Online Reinforcement Learning
- Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
- Enter the Void - Planning to Seek Entropy When Reward is Scarce
- Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approach
- Sequential Monte Carlo for Policy Optimization in Continuous POMDPs
- ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
- A Temporal Difference Method for Stochastic Continuous Dynamics
- Generative AI for Autonomous Driving: A Review
- RLVR-World: Training World Models with Reinforcement Learning
- Vid2World: Crafting Video Diffusion Models to Interactive World Models
- Policy-Driven World Model Adaptation for Robust Offline Model-based Reinforcement Learning
- Building spatial world models from sparse transitional episodic memories
- TD-GRPC: Temporal Difference Learning with Group Relative Policy Constraint for Humanoid Locomotion
- Zero-Shot Visual Generalization in Robot Manipulation
- EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
- Evolution imposes an inductive bias that alters and accelerates learning dynamics
- FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation
- Modeling Unseen Environments with Language-guided Composable Causal Components in Reinforcement Learning
- LaDi-WM: A Latent Diffusion-based World Model for Predictive Manipulation
- Generalization in Monitored Markov Decision Processes (Mon-MDPs)
- Video-Enhanced Offline Reinforcement Learning: A Model-Based Approach
- Pretraining a Shared Q-Network for Data-Efficient Offline Reinforcement Learning
- Self-Consistent Models and Values
- Reinforcement Learning with Latent Flow
- Coupled Distributional Random Expert Distillation for World Model Online Imitation Learning
- World Model-Based Learning for Long-Term Age of Information Minimization in Vehicular Networks
- Enhancing Policy Learning with World-Action Model
- Envisioning the Future, One Step at a Time
- Quo Vadis, World Modeling?
- Fractional Transfer Learning for Deep Model-Based Reinforcement Learning
- Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
- Adversarial observations in probabilistic State-Space Models for robust Reinforcement Learning
- Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning
- A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI
- Learn Proportional Derivative Controllable Latent Space from Pixels
- Fast LeWorldModel
- Vision Mamba in Remote Sensing: A Comprehensive Survey of Techniques, Applications and Outlook
- Uncertainty-aware Latent Safety Filters for Avoiding Out-of-Distribution Failures
- Solaris: Building a Multiplayer Video World Model in Minecraft
- The Invisible Hand of Physics: When Video Diffusion Models Know More Than They Show
- Looped World Models
- Beyond Language Modeling: An Exploration of Multimodal Pretraining
- Next Embedding Prediction Makes World Models Stronger
- World Action Models are Zero-shot Policies
- Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning
- Learning Latent Action World Models In The Wild
- Runtime Monitoring of Perception-Based Autonomous Systems via Embedding Temporal Logic
- On Training in Imagination
- Hallucination in World Models is Predictable and Preventable
- WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation
- Self-Monitoring Benefits from Structural Integration: Lessons from Metacognition in Continuous-Time Multi-Timescale Agents
- Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models
- From Kepler to Newton: Inductive Biases Guide Learned World Models in Transformers
- Nightmare Dreamer: Dreaming About Unsafe States And Planning Ahead
- Why LLMs Aren't Scientists Yet: Lessons from Four Autonomous Research Attempts
- Learned Perceptive Forward Dynamics Model for Safe and Platform-aware Robotic Navigation
- Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
- World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
- Designing Digital Humans with Ambient Intelligence
- GameTalk: Training LLMs for Strategic Conversation
- Enactive Artificial Intelligence: A Decision-Centric Architecture for Complex Systems
- Reinforced Imitation Learning by Free Energy Principle
- ATLAS: Adaptive Topological Learning with Abstract Successors for Continual Learning
- Vision-Language Model Predictive Control for Manipulation Planning and Trajectory Generation
- Distributional Active Inference
- Offline Robotic World Model: Learning Robotic Policies without a Physics Simulator
- PIN-WM: Learning Physics-INformed World Models for Non-Prehensile Manipulation
- Latent Diffusion Planning for Imitation Learning
- AI in a vat: Fundamental limits of efficient world modelling for agent sandboxing and interpretability
- CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning
- Provable Representation Learning for Imitation with Contrastive Fourier Features
- EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance
- PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models
- Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control
- HERA: Historical Evidence Routing Adapter for Physical Prediction in Latent World Models
- Quantum-Structured World Models (QSWMs) for Predictive Latent Dynamics
- MInCo: Mitigating Information Conflicts in Distracted Visual Model-based Reinforcement Learning
- Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination
- Beyond Static Forecasting: Unleashing the Power of World Models for Mobile Traffic Extrapolation
- Next-Future: Sample-Efficient Policy Learning for Robotic-Arm Tasks
- Emergence of Goal-Directed Behaviors via Active Inference with Self-Prior
- A Clean Slate for Offline Reinforcement Learning
- ViMo: A Generative Visual GUI World Model for App Agents
- MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
- State Estimation Using Particle Filtering in Adaptive Machine Learning Methods: Integrating Q-Learning and NEAT Algorithms with Noisy Radar Measurements
- Neural Motion Simulator: Pushing the Limit of World Models in Reinforcement Learning
- Improving World Models using Deep Supervision with Linear Probes
Discussions
Related