Embodied AI: From LLMs to World Models
2025/09/24 by Feng, Tongtong, Wang, Xin, Jiang, Yu-Gang +1 · 8 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Robotics (cs.RO)
paper · doi:10.48550/arxiv.2509.20021
Abstract
Embodied Artificial Intelligence (AI) is an intelligent system paradigm for achieving Artificial General Intelligence (AGI), serving as the cornerstone for various applications and driving the evolution from cyberspace to physical systems. Recent breakthroughs in Large Language Models (LLMs) and World Models (WMs) have drawn significant attention for embodied AI. On the one hand, LLMs empower embodied AI via semantic reasoning and task decomposition, bringing high-level natural language instructions and low-level natural language actions into embodied cognition. On the other hand, WMs empower embodied AI by building internal representations and future predictions of the external world, facilitating physical law-compliant embodied interactions. As such, this paper comprehensively explores the literature in embodied AI from basics to advances, covering both LLM driven and WM driven works. In particular, we first present the history, key technologies, key components, and hardware systems of embodied AI, as well as discuss its development via looking from unimodal to multimodal angle. We then scrutinize the two burgeoning fields of embodied AI, i.e., embodied AI with LLMs/multimodal LLMs (MLLMs) and embodied AI with WMs, meticulously delineating their indispensable roles in end-to-end embodied cognition and physical laws-driven embodied interactions. Building upon the above advances, we further share our insights on the necessity of the joint MLLM-WM driven embodied AI architecture, shedding light on its profound significance in enabling complex tasks within physical worlds. In addition, we examine representative applications of embodied AI, demonstrating its wide applicability in real-world scenarios. Last but not least, we point out future research directions of embodied AI that deserve further investigation.
Citations
- DETACH: Cross-domain Learning for Long-Horizon Tasks via Mixture of Disentangled Experts
- Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning
- Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning
- Embodied World Models Emerge from Navigational Task in Open-Ended Environments
- Modular-Cam: Modular Dynamic Camera-view Video Generation with LLM
- REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation
- Qwen2.5-Omni Technical Report
- Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation
- WorldModelBench: Judging Video Generation Models As World Models
- VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
- Magma: A Foundation Model for Multimodal AI Agents
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
- Emotional Face-to-Speech
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark
- JAQ: Joint Efficient Architecture Design and Low-Bit Quantization with Hardware-Software Co-Exploration
- Scale-adaptive UAV Geo-localization via Height-aware Partition Learning
- Understanding World or Predicting Future? A Comprehensive Survey of World Models
- Semantic Enhancement for Object SLAM with Heterogeneous Multimodal Large Language Model Agents
- π0: A Vision-Language-Action Flow Model for General Robot Control
- GPT-4o System Card
- Multimodal LLM Guided Exploration and Active Mapping using Fisher Information
- ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models
- MLLM as Retriever: Interactively Learning Multimodal Retrieval for Embodied Agents
- LaMMA-P: Generalizable Multi-Agent Long-Horizon Task Allocation and Planning with LM-Driven PDDL Planner
- TCDformer-based Momentum Transfer Model for Long-term Sports Prediction
- Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
- ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
- Meta-UAD: A Meta-Learning Scheme for User-level Network Traffic Anomaly Detection
- Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation
- Detecting Wildfires on UAVs with Real-time Segmentation Trained by Larger Teacher Models
- Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks
- Generative Multi-Agent Collaboration in Embodied AI: A Systematic Review
- LLaMAR: Long-Horizon Planning for Multi-Agent Robots in Partially Observable Environments
- Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models
- MFE-ETP: A Comprehensive Evaluation Benchmark for Multi-modal Foundation Models on Embodied Task Planning
- GenRL: Multimodal-foundation world models for generalization in embodied agents
- OpenVLA: An Open-Source Vision-Language-Action Model
- Hierarchical World Models as Visual Whole-Body Humanoid Controllers
- A Survey on Vision-Language-Action Models for Embodied AI
- Towards Efficient LLM Grounding for Embodied Multi-Agent Collaboration
- DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control
- Octo: An Open-Source Generalist Robot Policy
- Diffusion for World Modeling: Visual Details Matter in Atari
- QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
- Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond
- OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
- Large Language Models for UAVs: Current State and Pathways to the Future
- ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling
- A Survey on the Memory Mechanism of Large Language Model based Agents
- COMBO: Compositional World Models for Embodied Multi-Agent Cooperation
- HUGS: Holistic Urban 3D Scene Understanding via Gaussian Splatting
- LLM3:Large Language Model-based Task and Motion Planning with Motion Failure Reasoning
- Sora as a World Model? A Complete Survey on Text-to-Video Generation
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Genie: Generative Interactive Environments
- Revisiting Feature Prediction for Learning Visual Representations from Video
- AgentLens: Visual Analysis for Agent Behaviors in LLM-based Autonomous Systems
- AED: Adaptable Error Detection for Few-shot Imitation Policy
- The Essential Role of Causality in Foundation World Models for Embodied AI
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- A Survey on Robotics with Foundation Models: toward Embodied AI
- Executable Code Actions Elicit Better LLM Agents
- Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
- WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
- MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World
- AutoAct: Automatic Agent Learning from Scratch for QA via Self-Planning
- Gemini: Mapping and Architecture Co-exploration for Large-scale DNN Chiplet Accelerators
- EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI
- ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- CogAgent: A Visual Language Model for GUI Agents
- Foundation Models in Robotics: Applications, Challenges, and the Future
- MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active Perception
- VTimeLLM: Empower LLM to Grasp Video Moments
- GraphDreamer: Compositional 3D Scene Synthesis from Scene Graphs
- A-JEPA: Joint-Embedding Predictive Architecture Can Listen
- GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting
- Active Reasoning in an Open-World Environment
- RoboVQA: Multimodal Long-Horizon Reasoning for Robotics
- Qwen Technical Report
- AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- MC-JEPA: A Joint-Embedding Predictive Architecture for Self-Supervised Learning of Motion and Content Features
- Building Cooperative Embodied Agents Modularly with Large Language Models
- Embodied Task Planning with Large Language Models
- Physion++: Evaluating Physical Scene Understanding that Requires Online Inference of Different Physical Properties
- REFLECT: Summarizing Robot Experiences for Failure Explanation and Correction
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- Towards Label-free Scene Understanding by Vision Foundation Models
- Egocentric Planning for Scalable Embodied Task Achievement
- MADiff: Offline Multi-agent Learning with Diffusion Models
- AdaPlanner: Adaptive Planning from Feedback with Language Models
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought
- Language Models Meet World Models: Embodied Experiences Enhance Language Models
- DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation
- Visual Instruction Tuning
- ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments
- Micrograph segmentations for DDEVD
- Segment Anything
- RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding
- Reflexion: Language Agents with Verbal Reinforcement Learning
- GPT-4 Technical Report
- Transformer-based World Models Are Happy With 100k Interactions
- PaLM-E: An Embodied Multimodal Language Model
- Language Is Not All You Need: Aligning Perception with Language Models
- LLaMA: Open and Efficient Foundation Language Models
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- Mastering Diverse Domains through World Models
- RT-1: Robotics Transformer for Real-World Control at Scale
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models
- OpenScene: 3D Scene Understanding with Open Vocabularies
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
- Code as Policies: Language Model Programs for Embodied Control
- Transformers are Sample-Efficient World Models
- GAUDI: A Neural Architect for Immersive 3D Scene Generation
- DayDreamer: World Models for Physical Robot Learning
- MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge
- Multi-Agent Reinforcement Learning is a Sequence Modeling Problem
- A Generalist Agent
- Flamingo: a Visual Language Model for Few-Shot Learning
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- TrajGen: Generating Realistic and Diverse Trajectories with Reactive and Feasible Agent Behaviors for Autonomous Driving
- Diffusion Probabilistic Modeling for Video Generation
- Training language models to follow instructions with human feedback
- DreamingV2: Reinforcement Learning with Discrete World Models without Reconstruction
- TransDreamer: Reinforcement Learning with Transformer World Models
- SEAL: Self-supervised Embodied Active Learning using Exploration and 3D Consistency
- Masked Autoencoders Are Scalable Vision Learners
- TEACh: Task-driven Embodied Agents that Chat
- Language Grounding with 3D Objects
- A Survey of Embodied AI: From Simulators to Research Tasks
- Behavior From the Void: Unsupervised Active Pre-Training
- Learning Transferable Visual Models From Natural Language Supervision
- TrafficSim: Learning to Simulate Realistic Multi-Agent Behaviors
- GPT-3: Its Nature, Scope, Limits, and Consequences
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Mastering Atari with Discrete World Models
- EGO-Planner: An ESDF-free Gradient-based Local Planner for Quadrotors
- QPLEX: Duplex Dueling Multi-Agent Q-Learning
- Denoising Diffusion Probabilistic Models
- Learning to Explore using Active Neural SLAM
- MEMO: A Deep Network for Flexible Combination of Episodic Memories
- Spatial-Temporal Transformer Networks for Traffic Flow Forecasting
- Dream to Control: Learning Behaviors by Latent Imagination
- Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via\n Full-Stack Integration
- Are We Ready for Service Robots? The OpenLORIS-Scene Datasets for Lifelong SLAM
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- QTRAN: Learning to Factorize with Transformation for Cooperative\n Multi-Agent Reinforcement Learning
- A Short Survey On Memory Based Reinforcement Learning
- Habitat: A Platform for Embodied AI Research
- HAQ: Hardware-Aware Automated Quantization with Mixed Precision
- Learning Latent Dynamics for Planning from Pixels
- Model-Based Active Exploration
- World Models
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Memory Augmented Control Networks
- Learning to Look Around: Intelligently Exploring Unseen Environments for\n Unknown Tasks
- Proximal Policy Optimization Algorithms
- Attention Is All You Need
- AirSim: High-Fidelity Visual and Physical Simulation for Autonomous Vehicles
- In-Datacenter Performance Analysis of a Tensor Processing Unit
- Neural Episodic Control
- Generative Adversarial Imitation Learning
- Deep Residual Learning for Image Recognition
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- End-to-End Training of Deep Visuomotor Policies
- Playing Atari with Deep Reinforcement Learning
- ImageNet classification with deep convolutional neural networks
- Intelligence without representation
- I.—COMPUTING MACHINERY AND INTELLIGENCE
- EvolvingAgent: Curriculum Self-evolving Agent with Continual World Model for Long-Horizon Tasks
- U2UData+: A Scalable Swarm UAVs Autonomous Flight Dataset for Embodied Long-horizon Tasks
- Towards Rationality in Language and Multimodal Agents: A Survey
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Cited by
Related