Recurrent World Models Facilitate Policy Evolution
2018/09/04 by David Ha, Jürgen Schmidhuber, Ha, David +1 · 188 citations
Social Sciences · #FOS: Computer and information sciences #International Development and Aid #Machine Learning (cs.LG) #Machine Learning (stat.ML)
paper · pdf · doi:10.48550/arxiv.1809.01999
openalex publication_date 2018/09/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
A generative recurrent neural network is quickly trained in an unsupervised manner to model popular reinforcement learning environments through compressed spatio-temporal representations. The world model's extracted features are fed into compact and simple policies trained by evolution, achieving state of the art results in various environments. We also train our agent entirely inside of an environment generated by its own internal world model, and transfer this policy back into the actual environment. Interactive version of paper at https://worldmodels.github.io
Cited by
- Web World Models
- N0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation
- The Semantic Least-Energy Principle: A Hypothesis for Intelligence
- False Prophets: On the Security of World Models in Agentic Systems
- Dreamcrafter: Immersive Editing of 3D Radiance Fields Through Flexible, Generative Inputs and Outputs
- The Mental World of Large Language Models in Recommendation: A Benchmark on Association, Personalization, and Knowledgeability
- Beyond Sliding Windows: Learning to Manage Memory in Non-Markovian Environments
- InDRiVE: Reward-Free World-Model Pretraining for Autonomous Driving via Latent Disagreement
- Dexterous World Models
- Vidarc: Embodied Video Diffusion Model for Closed-loop Control
- Spectral Representation-based Reinforcement Learning
- World Models Unlock Optimal Foraging Strategies in Reinforcement Learning Agents
- VFMF: World Modeling by Forecasting Vision Foundation Model Features
- CERNet: Class-Embedding Predictive-Coding RNN for Unified Robot Motion, Recognition, and Confidence Estimation
- Agency at the Interface: Distinguishing Teleological from Structural Self-Organization via Internal Coarse-Graining and Downward Causation
- SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds
- VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory
- Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles
- Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
- Real-World Robot Control by Deep Active Inference With a Temporally Hierarchical World Model
- World Model Robustness via Surprise Recognition
- Learning Causal States Under Partial Observability and Perturbation
- How to Train Your Latent Control Barrier Function: Smooth Safety Filtering Under Hard-to-Model Constraints
- Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?
- Beyond Generative AI: World Models for Clinical Prediction, Counterfactuals, and Planning
- Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
- Dynamic Sparsity: Challenging Common Sparsity Assumptions for Learning World Models in Robotic Reinforcement Learning Benchmarks
- LLM-as-a-Judge: Toward World Models for Slate Recommendation Systems
- Enhancing Q-Value Updates in Deep Q-Learning via Successor-State Prediction
- Learning Interactive World Model for Object-Centric Reinforcement Learning
- Constructing the Umwelt: Cognitive Planning through Belief-Intent Co-Evolution
- Bootstrap Latent-Predictive Representations for Multitask Reinforcement Learning
- A Self-Supervised Auxiliary Loss for Deep RL in Partially Observable\n Settings
- Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method
- Social Attention for Autonomous Decision-Making in Dense Traffic
- Privacy-Preserving and Efficient Data Collection Scheme for AMI Networks Using Deep Learning
- Exploring Model-based Planning with Policy Networks
- DreamerPro: Reconstruction-Free Model-Based Reinforcement Learning with Prototypical Representations
- What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?
- Human Machine Social Hybrid Intelligence:A Collaborative Decision Making Framework for Large Model Agent Groups and Human Experts
- Deep Active Inference with Diffusion Policy and Multiple Timescale World Model for Real-World Exploration and Navigation
- I2-NeRF: Learning Neural Radiance Fields Under Physically-Grounded Media Interactions
- Predictive Coding Enhances Meta-RL To Achieve Interpretable Bayes-Optimal Belief Representation Under Partial Observability
- Evaluating Video Models as Simulators of Multi-Person Pedestrian Trajectories
- LyTimeT: Towards Robust and Interpretable State-Variable Discovery
- Contrastive Variational Reinforcement Learning for Complex Observations
- Socialized Learning and Emergent Behaviors in Multi-Agent Systems based on Multimodal Large Language Models
- Efficient Model-Based Reinforcement Learning for Robot Control via Online Optimization
- A Comprehensive Survey on World Models for Embodied AI
- XLVIN: eXecuted Latent Value Iteration Nets
- Cosmos-Surg-dVRK: World Foundation Model-based Automated Online Evaluation of Surgical Robot Policy Learning
- SkyDreamer: Interpretable End-to-End Vision-Based Drone Racing with Model-Based Reinforcement Learning
- Ego-Vision World Model for Humanoid Contact Planning
- Unsupervised State Representation Learning in Atari
- SoundReactor: Frame-level Online Video-to-Audio Generation
- Improving Generalization in Meta Reinforcement Learning using Learned Objectives
- Towards Neurocognitive-Inspired Intelligence: From AI's Structural Mimicry to Human-Like Functional Cognition
- Agent Learning via Early Experience
- Generative World Modelling for Humanoids: 1X World Model Challenge Technical Report
- Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models
- DREAMer-VXS: A Latent World Model for Sample-Efficient AGV Exploration in Stochastic, Unobserved Environments
- On the role of planning in model-based deep reinforcement learning
- Model-based micro-data reinforcement learning: what are the crucial model properties and which model to choose?
- Model-Based Episodic Memory Induces Dynamic Hybrid Controls
- Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft
- World Model for AI Autonomous Navigation in Mechanical Thrombectomy
- Allocentric flocking
- Model-Based Reinforcement Learning for Atari
- Dual-Scale World Models for LLM Agents Towards Hard-Exploration Problems
- Adversarial Diffusion for Robust Reinforcement Learning
- Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation
- Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
- VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
- Progressive growing of self-organized hierarchical representations for exploration
- Context and Diversity Matter: The Emergence of In-Context Learning in World Models
- Model-Based Reinforcement Learning under Random Observation Delays
- From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery
- System 0/1/2/3: Quad-Process Theory for Multitimescale Embodied Collective Cognitive Systems
- Understanding the hippocampus as an apex of the cortical hierarchy and self-supervised predictive learning engine
- World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation
- Fully Learnable Neural Reward Machines
- On the model-based stochastic value gradient for continuous reinforcement learning
- Remote Sensing-Oriented World Model
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
- SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models
- Foundation Models as World Models: A Foundational Study in Text-Based GridWorlds
- Modeling Worlds in Text
- Explainable Autonomous Robots: A Survey and Perspective
- Training Larger Networks for Deep Reinforcement Learning
- From Next Token Prediction to (STRIPS) World Models -- Preliminary Results
- Continuous 3D Multi-Channel Sign Language Production via Progressive Transformers and Mixture Density Networks
- Learning Knowledge Graph-based World Models of Textual Environments
- Avoiding Knowledge Edit Skipping in Multi-hop Question Answering with Guided Decomposition
- Language (Re)modelling: Towards Embodied Language Understanding
- Spiking Neural Networks for Continuous Control via End-to-End Model-Based Learning
- Long-Horizon Visual Planning with Goal-Conditioned Hierarchical Predictors
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
- Improving Generative Imagination in Object-Centric World Models
- Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
- Continual State Representation Learning for Reinforcement Learning using Generative Replay
- Vision-Based Autonomous Car Racing Using Deep Imitative Reinforcement Learning
- Data-Efficient Reinforcement Learning for Malaria Control
- Goal-Directed Planning by Reinforcement Learning and Active Inference
- MAPF-World: Action World Model for Multi-Agent Path Finding
- Finite Automata Extraction: Low-data World Model Learning as Programs from Gameplay Video
- Self-organization of action hierarchy and compositionality by reinforcement learning with recurrent neural networks
- Meta-Reinforcement Learning for Adaptive Motor Control in Changing Robot Dynamics and Environments
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
- How Safe Will I Be Given What I Saw? Calibrated Prediction of Safety Chances for Image-Controlled Autonomy
- Causal World Models by Unsupervised Deconfounding of Physical Dynamics
- Channel Decomposition into Painting Actions
- Bounding Distributional Shifts in World Modeling through Novelty Detection
- In-Context Reinforcement Learning via Communicative World Models
- ReSim: Reliable World Simulation for Autonomous Driving
- Dynamics-Aware Unsupervised Discovery of Skills
- CoEx -- Co-evolving World-model and Exploration
- Learning to Paint With Model-based Deep Reinforcement Learning
- Distilling a Hierarchical Policy for Planning and Control via Representation and Reinforcement Learning
- A Roadmap for Robust End-to-End Alignment
- Back to the Features: DINO as a Foundation for Video World Models
- SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models
- Visual Grounding of Learned Physical Models
- Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models
- Variational Recurrent Models for Solving Partially Observable Control Tasks
- Continual Learning Using World Models for Pseudo-Rehearsal
- Iterative Model-Based Reinforcement Learning Using Simulations in the Differentiable Neural Computer
- Improving Gradient Estimation in Evolutionary Strategies With Past Descent Directions
- A Perspective on Objects and Systematic Generalization in Model-Based RL
- Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories
- Never Stop Learning: The Effectiveness of Fine-Tuning in Robotic Reinforcement Learning
- Variational State-Space Models for Localisation and Dense 3D Mapping in 6 DoF
- Domain-Adversarial and Conditional State Space Model for Imitation Learning
- SketchTransfer: A Challenging New Task for Exploring Detail-Invariance and the Abstractions Learned by Deep Networks
- World Model-Based End-to-End Scene Generation for Accident Anticipation in Autonomous Driving
- Deep Active Inference as Variational Policy Gradients
- Learning to Predict Without Looking Ahead: World Models Without Forward Prediction
- Deep Learning: Our Miraculous Year 1990-1991
- Learning Belief Representations for Imitation Learning in POMDPs
- Reinforcement Learning-based Visual Navigation with Information-Theoretic Regularization
- Graph World Model
- GenAI-based Multi-Agent Reinforcement Learning towards Distributed Agent Intelligence: A Generative-RL Agent Perspective
- NeuralOS: Towards Simulating Operating Systems via Neural Generative Models
- AquaJEPA: Action-Conditioned Multimodal Predictive Representations for Underwater Robot Dynamics
- Spatial-Temporal Aware Visuomotor Diffusion Policy Learning
- Latent Actions from Factorized Transition Effects under Agent Ambiguity
- Forward Prediction for Physical Reasoning
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
- Accurate and Efficient World Modeling with Masked Latent Transformers
- Dyn-O: Building Structured World Models with Object-Centric Representations
- Mitigating Goal Misgeneralization via Minimax Regret
- ActionParty: Multi-Subject Action Binding in Generative Video Games
- NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments
- R-SQAIR: Relational Sequential Attend, Infer, Repeat
- Mathematical Reasoning in Latent Space
- RoboScape: Physics-informed Embodied World Model
- SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model
- Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
- Koopman Q-learning: Offline Reinforcement Learning via Symmetries of\n Dynamics
- From 2D to 3D Cognition: A Brief Survey of General World Models
- Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline Environment
- ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model
- Adapting Vision-Language Models for Evaluating World Models
- Reimagination with Test-time Observation Interventions: Distractor-Robust World Model Predictions for Visual Model Predictive Control
- Measuring Intent Comprehension in LLMs
- Interpretability and Generalization Bounds for Learning Spatial Physics
- GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation
- Foundation Models in Autonomous Driving: A Survey on Scenario Generation and Scenario Analysis
- Detection of False-Reading Attacks in the AMI Net-Metering System
- TARDIS STRIDE: A Spatio-Temporal Road Image Dataset and World Model for Autonomy
- Planning from Pixels using Inverse Dynamics Models
- Bootstrapping World Models from Dynamics Models in Multimodal Foundation Models
- Reinforcement Learning for Robotics and Control with Active Uncertainty\n Reduction
- Video World Models with Long-term Spatial Memory
- Skill Transfer in Deep Reinforcement Learning under Morphological Heterogeneity
- Regularizing Model-Based Planning with Energy-Based Models
- Simple, Good, Fast: Self-Supervised World Models Free of Baggage
- Muesli: Combining Improvements in Policy Optimization
- Learning Abstract World Models with a Group-Structured Latent Space
- Continuous Control for Searching and Planning with a Learned Model
- Optimizing Sensory Neurons: Nonlinear Attention Mechanisms for Accelerated Convergence in Permutation-Invariant Neural Networks for Reinforcement Learning
- Safe Reinforcement Learning with Mixture Density Network: A Case Study in Autonomous Highway Driving
- Toward Memory-Aided World Models: Benchmarking via Spatial Consistency
- Efficient Reinforcement Learning for StarCraft by Abstract Forward Models and Transfer Learning
- VRAG: Learning World Models for Interactive Video Generation
- StateSpaceDiffuser: Bringing Long Context to Diffusion World Models
Related