A Survey: Learning Embodied Intelligence from Physical Simulators and World Models
2025/07/01 by Xiaoxiao Long, Qingrui Zhao, Long, Xiaoxiao +33 · 1 voice · 18 citations
Computer Science · Engineering · Psychology · #Action Observation and Synchronization #FOS: Computer and information sciences #Human Motion and Animation #Robotics (cs.RO) #Social Robot Interaction and HRI #cs.RO
paper · pdf · doi:10.48550/arxiv.2507.00917
openalex publication_date 2025/07/01 · arxiv published 2025/07/01 · arxiv updated 2025/09/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The pursuit of artificial general intelligence (AGI) has placed embodied intelligence at the forefront of robotics research. Embodied intelligence focuses on agents capable of perceiving, reasoning, and acting within the physical world. Achieving robust embodied intelligence requires not only advanced perception and control, but also the ability to ground abstract cognition in real-world interactions. Two foundational technologies, physical simulators and world models, have emerged as critical enablers in this quest. Physical simulators provide controlled, high-fidelity environments for training and evaluating robotic agents, allowing safe and efficient development of complex behaviors. In contrast, world models empower robots with internal representations of their surroundings, enabling predictive planning and adaptive decision-making beyond direct sensory input. This survey systematically reviews recent advances in learning embodied AI through the integration of physical simulators and world models. We analyze their complementary roles in enhancing autonomy, adaptability, and generalization in intelligent robots, and discuss the interplay between external simulation and internal modeling in bridging the gap between simulated training and real-world deployment. By synthesizing current progress and identifying open challenges, this survey aims to provide a comprehensive perspective on the path toward more capable and generalizable embodied AI systems. We also maintain an active repository that contains up-to-date literature and open-source projects at https://github.com/NJU3DV-LoongGroup/Embodied-World-Models-Survey.
Citations
- Dynamics-Aligned Latent Imagination in Contextual World Models for Zero-Shot Generalization
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
- Imaginative World Modeling with Scene Graphs for Embodied Agent Navigation
- IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model
- EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling
- Epona: Autoregressive Diffusion World Model for Autonomous Driving
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- A Unified and General Humanoid Whole-Body Controller for Versatile Locomotion
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
- Humanoid World Models: Open World Foundation Models for Humanoid Robotics
- DeepVerse: 4D Autoregressive Video Generation as a World Model
- DexMachina: Functional Retargeting for Bimanual Dexterous Manipulation
- GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control
- OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory Generation
- WorldEval: World Model as Real-World Robot Policies Evaluator
- Vid2World: Crafting Video Diffusion Models to Interactive World Models
- FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation
- EnerVerse-AC: Envisioning Embodied Environments with Action Condition
- Vision-Language-Action (VLA) Models: Concepts, Progress, Applications and Challenges
- Visual Imitation Enables Contextual Humanoid Control
- Learning 3D Persistent Embodied World Models
- TesserAct: Learning 4D Embodied World Models
- PIN-WM: Learning Physics-INformed World Models for Non-Prehensile Manipulation
- Offline Robotic World Model: Learning Robotic Policies without a Physics Simulator
- Neural Motion Simulator: Pushing the Limit of World Models in Reinforcement Learning
- Dexterous Manipulation through Imitation Learning: A Survey
- SocialGesture: Delving into Multi-person Gesture Understanding
- End-to-End Driving with Online Trajectory Evaluation via BEV World Model
- ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning
- GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
- Aether: Geometric-Aware Unified World Modeling
- AdaWorld: Learning Adaptable World Models with Latent Actions
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
- Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control
- UniGoal: Towards Universal Zero-shot Goal-oriented Navigation
- LUMOS: Language-Conditioned Imitation Learning with World Models
- DexGrasp Anything: Towards Universal Robotic Dexterous Grasping with Physics Awareness
- BEHAVIOR Robot Suite: Streamlining Real-World Whole-Body Manipulation for Everyday Household Activities
- DualDiff+: Dual-Branch Diffusion for High-Fidelity Video Generation with Reward Guidance
- DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping
- HiFAR: Multi-Stage Curriculum Learning for High-Dynamics Humanoid Fall Recovery
- Accelerating Model-Based Reinforcement Learning with State-Space World Models
- Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
- Learning Humanoid Locomotion with World Model Reconstruction
- Magma: A Foundation Model for Multimodal AI Agents
- Learning Humanoid Standing-up Control across Diverse Postures
- ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills
- Embrace Collisions: Humanoid Shadowing for Deployable Contact-Agnostics Motions
- HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
- AdaWM: Adaptive World Model based Planning for Autonomous Driving
- Robotic World Model: A Neural Network Simulator for Robust Policy Optimization in Robotics
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- Self-Supervised Representation Learning with Joint Embedding Predictive Architecture for Automotive LiDAR Object Detection
- MobileH2R: Learning Generalizable Human to Mobile Robot Handover Exclusively from Scalable and Diverse Synthetic Data
- Cosmos World Foundation Model Platform for Physical AI
- EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation
- DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT
- DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers
- LLaVA-SLT: Visual Language Tuning for Sign Language Translation
- Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination
- An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training
- ExBody2: Advanced Expressive Humanoid Whole-Body Control
- GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
- Video Representation Learning with Joint-Embedding Predictive Architectures
- GaussianWorld: Gaussian World Model for Streaming 3D Occupancy\n Prediction
- Doe-1: Closed-Loop Autonomous Driving with Large World Model
- CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs
- Pysical Informed Driving World Model
- FLIP: Flow-Centric Generative Planning as General-Purpose Manipulation World Model
- STIV: Scalable Text and Image Conditioned Video Generation
- ACT-Bench: Towards Action Controllable World Models for Autonomous Driving
- Navigation World Models
- HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving
- InfinityDrive: Breaking Time Limits in Driving World Models
- ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration
- DextrAH-RGB: Visuomotor Policies to Grasp Anything with Dexterous Hands
- Understanding World or Predicting Future? A Comprehensive Survey of World Models
- MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control
- Learning Humanoid Locomotion with Perceptive Internal Model
- WHALE: Towards Generalizable and Scalable World Models for Embodied Decision-making
- π0: A Vision-Language-Action Flow Model for General Robot Control
- Neural Attention Field: Emerging Point Relevance in 3D Scenes for One-Shot Dexterous Grasping
- DexGraspNet 2.0: Learning Generative Dexterous Grasping in Large-scale Synthetic Cluttered Scenes
- WorldSimBench: Towards Video Generation Models as World Simulators
- DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation
- ALOHA Unleashed: A Simple Recipe for Robot Dexterity
- Harmon: Whole-Body Motion Generation of Humanoid Robots from Language Descriptions
- OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation
- PIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic Manipulation
- DOME: Taming Diffusion Model into High-Fidelity Controllable Occupancy World Model
- RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
- SG-Nav: Online 3D Scene Graph Prompting for LLM-based Zero-shot Object Navigation
- GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
- D(R,O) Grasp: A Unified Representation of Robot and Object Interaction for Cross-Embodiment Dexterous Grasping
- ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI
- World Model-based Perception for Visual Legged Locomotion
- Mitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World Models
- TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
- RenderWorld: World Model with Self-Supervised 3D Label
- MRAC Track 1: 2nd Workshop on Multimodal, Generative and Responsible\n Affective Computing
- Towards Social AI: A Survey on Understanding Social Interactions
- OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving
- Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving
- DriveArena: A Closed-loop Generative Simulation Platform for Autonomous Driving
- CarFormer: Self-Driving with Learned Object-Centric Representations
- Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning
- Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
- DextrAH-G: Pixels-to-Action Dexterous Arm-Hand Grasping with Geometric Fabrics
- Open-TeleVision: Teleoperation with Immersive Active Visual Feedback
- From Efficient Multimodal Models to World Models: A Survey
- NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking
- Humanoid Parkour Learning
- HumanPlus: Humanoid Shadowing and Imitation from Humans
- OpenVLA: An Open-Source Vision-Language-Action Model
- OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning
- Enhancing End-to-End Autonomous Driving with Latent World Model
- UnO: Unsupervised Occupancy Fields for Perception and Forecasting
- Learning-based legged locomotion; state of the art and future perspectives
- OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving
- Grasp as You Say: Language-guided Dexterous Grasp Generation
- Hierarchical World Models as Visual Whole-Body Humanoid Controllers
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability
- iVideoGPT: Interactive VideoGPTs are Scalable World Models
- MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes
- CDM-MPC: An Integrated Dynamic Planning and Control Framework for Bipedal Robots Jumping
- DriveWorld: 4D Pre-trained Scene Understanding via World Models for Autonomous Driving
- Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond
- The Call for Socially Aware Language Technologies
- DiffuseLoco: Real-Time Legged Locomotion Control with Diffusion from Offline Datasets
- SpringGrasp: Synthesizing Compliant, Dexterous Grasps under Shape Uncertainty
- RoboDreamer: Learning Compositional World Models for Robot Imagination
- JRDB-Social: A Multifaceted Robotic Dataset for Understanding of Context and Dynamics of Human Interactions Within Social Groups
- LidarDM: Generative LiDAR Simulation in a Generated World
- TriHelper: Zero-Shot Object Navigation with Dynamic Assistance
- 3D-VLA: A 3D Vision-Language-Action Generative World Model
- GenAD: Generalized Predictive Model for Autonomous Driving
- DexCap: Scalable and Portable Mocap Data Collection System for Dexterous Manipulation
- DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation
- Towards learning-based planning:The nuPlan benchmark for real-world autonomous driving
- 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations
- Whole-body Humanoid Robot Locomotion with Human Reference
- Expressive Whole-Body Control for Humanoid Robots
- Genie: Generative Interactive Environments
- Revisiting Feature Prediction for Learning Visual Representations from Video
- Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots
- Reinforcement Learning for Versatile, Dynamic, and Robust Bipedal Locomotion Control
- Adaptive Mobile Manipulation for Articulated Objects In the Open World
- Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation
- GenH2R: Learning Generalizable Human-to-Robot Handover via Scalable Simulation, Demonstration, and Imitation
- Visual Point Cloud Forecasting enables Scalable Autonomous Driving
- VideoPoet: A Large Language Model for Zero-Shot Video Generation
- The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective
- A Survey on Robotic Manipulation of Deformable Objects: Recent Advances, Open Challenges and New Frontiers
- WoVoGen: World Volume-aware Diffusion for Controllable Multi-camera Driving Scene Generation
- Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
- Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving
- Panacea: Panoramic and Controllable Video Generation for Autonomous Driving
- OccWorld: Learning a 3D Occupancy World Model for Autonomous Driving
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- ADriver-I: A General World Model for Autonomous Driving
- TWIST: Teacher-Student World Model Distillation for Efficient Sim-to-Real Transfer
- Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion
- SparseDFF: Sparse-View Feature Distillation for One-Shot Dexterous Manipulation
- Finetuning Offline World Models in the Real World
- MagicDrive: Street View Generation with Diverse 3D Geometry Control
- HarmonyDream: Task Harmonization Inside World Models
- GAIA-1: A Generative World Model for Autonomous Driving
- Learning Vision-Based Bipedal Locomotion for Challenging Terrain
- MoDem-V2: Visuo-Motor World Models for Real-World Robot Manipulation
- A Survey of Imitation Learning: Algorithms, Recent Developments, and Challenges
- Stabilize to Act: Learning to Coordinate for Bimanual Manipulation
- UniWorld: Autonomous Driving Pre-training via World Models
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- SafeDreamer: Safe Reinforcement Learning with World Models
- Structured World Models from Human Videos
- Surfer: Progressive Reasoning with World Models for Robotic Manipulation
- Neural LerPlane Representations for Fast 4D Reconstruction of Deformable Tissues
- Fast-Grasp'D: Dexterous Multi-finger Grasp Generation Through Differentiable Simulation
- Optimizing Bipedal Locomotion for The 100m Dash With Comparison to Human Running
- Pre-training Contextualized World Models with In-the-wild Videos for Reinforcement Learning
- NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario
- Video Prediction Models as Rewards for Reinforcement Learning
- Occ3D: A Large-Scale 3D Occupancy Prediction Benchmark for Autonomous Driving
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
- Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models
- Legs as Manipulator: Pushing Quadrupedal Agility Beyond Locomotion
- Rotating without Seeing: Towards In-hand Dexterity through Touch
- GPT-4 Technical Report
- Transformer-based World Models Are Happy With 100k Interactions
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
- Diffusion policy: Visuomotor policy learning via action diffusion
- TrafficBots: Towards World Models for Autonomous Driving Simulation and Motion Prediction
- PaLM-E: An Embodied Multimodal Language Model
- Open-World Object Manipulation using Pre-trained Vision-Language Models
- ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- Diffusion-based Generation, Optimization, and Planning in 3D Scenes
- Mastering Diverse Domains through World Models
- RT-1: Robotics Transformer for Real-World Control at Scale
- Learning Robust Real-World Dexterous Grasping Policies via Implicit Shape Augmentation
- Model-Based Imitation Learning for Urban Driving
- Enhance Sample Efficiency and Robustness of End-to-end Urban Autonomous Driving via Semantic Masked World Model
- Imagen Video: High Definition Video Generation with Diffusion Models
- Code as Policies: Language Model Programs for Embodied Control
- Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- DayDreamer: World Models for Physical Robot Learning
- CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
- Iso-Dream: Isolating and Leveraging Noncontrollable Visual Dynamics in World Models
- HandoverSim: A Simulation Framework and Benchmark for Human-to-Robot Object Handovers
- Video Diffusion Models
- DVGG: Deep Variational Grasp Generation for Dextrous Manipulation
- Adversarial Motion Priors Make Good Substitutes for Complex Reward Functions
- CODA: A Real-World Road Corner Case Dataset for Object Detection in Autonomous Driving
- DreamingV2: Reinforcement Learning with Discrete World Models without Reconstruction
- TransDreamer: Reinforcement Learning with Transformer World Models
- DexVIP: Learning Dexterous Grasping with Human Hand Pose Priors from Video
- Neural Descriptor Fields: SE(3)-Equivariant Object Representations for Manipulation
- NÜWA: Visual Synthesis Pre-training for Neural visUal World creAtion
- A System for General In-Hand Object Re-Orientation
- DreamerPro: Reconstruction-Free Model-Based Reinforcement Learning with Prototypical Representations
- Discovering and Achieving Goals via World Models
- Human-robot collaboration and machine learning: a systematic review of recent research
- Graph Based Network with Contextualized Representations of Turns in Dialogue
- Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
- Efficient Anytime CLF Reactive Planning System for a Bipedal Robot on Undulating Terrain
- Fast Contact-Implicit Model-Predictive Control
- Blind Bipedal Stair Traversal via Sim-to-Real Reinforcement Learning
- Hand-Object Contact Consistency Reasoning for Human Grasps Generation
- A Survey of Embodied AI: From Simulators to Research Tasks
- Where2Act: From Pixels to Actions for Articulated 3D Objects
- Adaptive Force-based Control for Legged Robots
- Mastering Atari with Discrete World Models
- Learning Dexterous Grasping with Object-Centric Visual Affordances
- One Thousand and One Hours: Self-driving Motion Prediction Dataset
- Planning to Explore via Self-Supervised World Models
- SAPIEN: A SimulAted Part-based Interactive ENvironment
- Deep Differentiable Grasp Planner for High-DOF Grippers
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
- 6-DOF Grasping for Target-driven Object Manipulation in Clutter
- Dream to Control: Learning Behaviors by Latent Imagination
- Cable Manipulation with a Tactile-Reactive Gripper
- Human-in-the-loop Robotic Manipulation Planning for Collaborative Assembly
- DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in Conversation
- A Survey of Autonomous Driving: Common Practices and Emerging Technologies
- Dynamic Walking with Compliance on a Cassie Bipedal Robot
- AMASS: Archive of Motion Capture as Surface Shapes
- nuScenes: A multimodal dataset for autonomous driving
- Iterative Reinforcement Learning Based Design of Dynamic Locomotion Skills for Cassie
- Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation
- Learning Latent Dynamics for Planning from Pixels
- Baidu Apollo EM Motion Planner
- Considering Human Behavior in Motion Planning for Smooth Human-Robot Collaboration in Close Proximity
- Bipedal Hopping: Reduced-order Model Embedding via Optimization-based Control
- World Models
- Neural Discrete Representation Learning
- PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes
- CARLA: An Open Urban Driving Simulator
- AffordanceNet: An End-to-End Deep Learning Approach for Object Affordance Detection
- Behavior Trees in Robotics and AI
- Emergence of Locomotion Behaviours in Rich Environments
- End-to-end Learning of Driving Models from Large-scale Video Datasets
- Vision meets robotics: The KITTI dataset
- Reinforcement learning in robotics: A survey
- Dual arm manipulation—A survey
- Sampling-based Algorithms for Optimal Motion Planning
- Dialogue Act Modeling for Automatic Tagging and Recognition of Conversational Speech
- RoboTransfer: Geometry-Consistent Video Diffusion for Robotic Visual Policy Transfer
- DreamGen: Unlocking Generalization in Robot Learning through Video World Models
- Imagine-2-Drive: Leveraging High-Fidelity World Models via Multi-Modal Diffusion Policies
- Generalizable Humanoid Manipulation with 3D Diffusion Policies
Cited by
Discussions
Related