A Survey on Vision-Language-Action Models for Embodied AI
2024/05/23 by Ma, Yueen, Song, Zixing, Zhuang, Yuzheng +2 · 1 voice · 81 citations
#Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics (cs.RO)
paper · doi:10.48550/arxiv.2405.14093
Abstract
Embodied AI is widely recognized as a key element of artificial general intelligence because it involves controlling embodied agents to perform tasks in the physical world. Building on the success of large language models and vision-language models, a new category of multimodal models -- referred to as vision-language-action models (VLAs) -- has emerged to address language-conditioned robotic tasks in embodied AI by leveraging their distinct ability to generate actions. In recent years, a myriad of VLAs have been developed, making it imperative to capture the rapidly evolving landscape through a comprehensive survey. To this end, we present the first survey on VLAs for embodied AI. This work provides a detailed taxonomy of VLAs, organized into three major lines of research. The first line focuses on individual components of VLAs. The second line is dedicated to developing control policies adept at predicting low-level actions. The third line comprises high-level task planners capable of decomposing long-horizon tasks into a sequence of subtasks, thereby guiding VLAs to follow more general user instructions. Furthermore, we provide an extensive summary of relevant resources, including datasets, simulators, and benchmarks. Finally, we discuss the challenges faced by VLAs and outline promising future directions in embodied AI. We have created a project associated with this survey, which is available at https://github.com/yueen-ma/Awesome-VLA.
Cited by
- ProGuard: Towards Proactive Multimodal Safeguard
- Agentic Physical AI toward a Domain-Specific Foundation Model for Energy Systems: A Case Study on Nuclear Reactor Control
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
- MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents
- Physical AI Governance: From Theory to Practice Across Life Cycle
- Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents
- VL4Gaze: Unleashing Vision-Language Models for Gaze Following
- Generative Human-Object Interaction Detection via Differentiable Cognitive Steering of Multi-modal LLMs
- CoDrone: Autonomous Drone Navigation Assisted by Edge and Cloud Foundation Models
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space
- From Human Intention to Action Prediction: A Comprehensive Benchmark for Intention-driven End-to-End Autonomous Driving
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models
- VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer
- Sigma: The Key for Vision-Language-Action Models toward Telepathic Alignment
- ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
- Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
- Eq.Bot: Enhance Robotic Manipulation Learning via Group Equivariant Canonicalization
- AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
- Learning by Neighbor-Aware Semantics, Deciding by Open-form Flows: Towards Robust Zero-Shot Skeleton Action Recognition
- SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
- 10 Open Challenges Steering the Future of Vision-Language-Action Models
- Systematizing LLM Persona Design: A Four-Quadrant Technical Taxonomy for AI Companion Applications
- DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
- Can Vision-Language-Action Models Learn from Real-World Data Continually without Forgetting?
- Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical World
- MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation
- RoboOmni: Proactive Robot Manipulation in Omni-modal Context
- A Survey on Efficient Vision-Language-Action Models
- NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?
- SafeCoop: Unravelling Full Stack Safety in Agentic Collaborative Driving
- Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots
- Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
- RoboHiMan: A Hierarchical Evaluation Paradigm for Compositional Generalization in Long-Horizon Manipulation
- B2N3D: Progressive Learning from Binary to N-ary Relationships for 3D Object Grounding
- Goal-oriented Backdoor Attack against Vision-Language-Action Models via Physical Objects
- Executable Analytic Concepts as the Missing Link Between VLM Insight and Precise Manipulation
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- RLinf-VLA: A Unified and Efficient Framework for VLA+RL Training
- MetaVLA: Unified Meta Co-training For Efficient Embodied Adaption
- HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
- VLA Model Post-Training via Action-Chunked PPO and Self Behavior Cloning
- OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation
- PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion
- VIVA+: Human-Centered Situational Decision-Making
- PhysiAgent: An Embodied Agent Framework in Physical World
- HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection
- Measuring Physical-World Privacy Awareness of Large Language Models: An Evaluation Benchmark
- Transferring Vision-Language-Action Models to Industry Applications: Architectures, Performance, and Challenges
- RoboView-Bias: Benchmarking Visual Bias in Embodied Agents for Robotic Manipulation
- MoWM: Mixture-of-World-Models for Embodied Planning via Latent-to-Pixel Feature Modulation
- Embodied AI: From LLMs to World Models
- Beyond Human Demonstrations: Diffusion-Based Reinforcement Learning to Generate Data for VLA Training
- Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems
- Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs
- HLG: Comprehensive 3D Room Construction via Hierarchical Layout Generation
- Eva-VLA: Evaluating Vision-Language-Action Models' Robustness Under Real-World Physical Variations
- Growing with Your Embodied Agent: A Human-in-the-Loop Lifelong Code Generation Framework for Long-Horizon Manipulation Skills
- LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA
- ComposableNav: Instruction-Following Navigation in Dynamic Environments via Composable Diffusion
- ByteWrist: A Parallel Robotic Wrist Enabling Flexible and Anthropomorphic Motion for Confined Spaces
- MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
- Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges
- Decoding RobKiNet: Insights into Efficient Training of Robotic Kinematics Informed Neural Network
- Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
- Robotic Manipulation Framework Based on Semantic Keypoints for Packing Shoes of Different Sizes, Shapes, and Softness
- DeepThink3D: Enhancing Large Language Models with Programmatic Reasoning in Complex 3D Situated Reasoning Tasks
- ContextualLVLM-Agent: A Holistic Framework for Multi-Turn Visually-Grounded Dialogue and Complex Instruction Following
- Survey of Vision-Language-Action Models for Embodied Manipulation
- Sensing, Social, and Motion Intelligence in Embodied Navigation: A Comprehensive Survey
- Multimodal Data Storage and Retrieval for Embodied AI: A Survey
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- Human Centric General Physical Intelligence for Agile Manufacturing Automation
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- Δ-AttnMask: Attention-Guided Masked Hidden States for Efficient Data Selection and Augmentation
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
- Learning Robust Intervention Representations with Delta Embeddings
- Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
Discussions
Related