Survey of Vision-Language-Action Models for Embodied Manipulation
2025/08/21 by Haoran Li, Li, Haoran, Yu‐Hui Chen +13 · 4 citations
Computer Science · Psychology · #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics #Social Robot Interaction and HRI
paper · pdf · doi:10.48550/arxiv.2508.15201
Abstract
Embodied intelligence systems, which enhance agent capabilities through continuous environment interactions, have garnered significant attention from both academia and industry. Vision-Language-Action models, inspired by advancements in large foundation models, serve as universal robotic control frameworks that substantially improve agent-environment interaction capabilities in embodied intelligence systems. This expansion has broadened application scenarios for embodied AI robots. This survey comprehensively reviews VLA models for embodied manipulation. Firstly, it chronicles the developmental trajectory of VLA architectures. Subsequently, we conduct a detailed analysis of current research across 5 critical dimensions: VLA model structures, training datasets, pre-training methods, post-training methods, and model evaluation. Finally, we synthesize key challenges in VLA development and real-world deployment, while outlining promising future research directions.
Citations
- RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim Translation
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- Igniting VLMs toward the Embodied Space
- Towards Human-level Intelligence via Human-like Whole-Body Manipulation
- Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
- GR-3 Technical Report
- EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
- Vision Language Action Models in Robotic Manipulation: A Systematic Review
- CL3R: 3D Reconstruction and Contrastive Learning for Enhanced Robotic Manipulation Representations
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
- A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
- WorldVLA: Towards Autoregressive Action World Model
- Parallels Between VLA Model Post-Training and Human Motor Learning: Progress, Challenges, and Trends
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
- RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
- ControlVLA: Few-shot Object-centric Adaptation for Pre-trained Vision-Language-Action Models
- SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model Acceleration
- From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
- EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models
- TGRPO :Fine-tuning Vision-Language-Action Model via Trajectory-wise Group Relative Policy Optimization
- Fast ECoT: Efficient Embodied Chain-of-Thought via Thoughts Reuse
- BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
- BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models
- Real-Time Execution of Action Chunking Flow Policies
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
- ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
- ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation
- Hume: Introducing System-2 Thinking in Visual-Language-Action Model
- What Can RL Bring to VLA Generalization? An Empirical Study
- WorldEval: World Model as Real-World Robot Policies Evaluator
- VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
- Interactive Post-Training for Vision-Language-Action Models
- Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization
- InSpire: Vision-Language-Action Models with Intrinsic Spatial Reasoning
- Vid2World: Crafting Video Diffusion Models to Interactive World Models
- Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation
- OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning
- VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation
- Training Strategies for Efficient Embodied Reasoning
- ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- Vision-Language-Action (VLA) Models: Concepts, Progress, Applications and Challenges
- GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data
- OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
- NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
- π0.5: a Vision-Language-Action Model with Open-World Generalization
- AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
- DyWA: Dynamics-adaptive World Action Model for Generalizable Non-prehensile Manipulation
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
- TLA: Tactile-Language-Action Model for Contact-Rich Manipulation
- FP3: A 3D Foundation Policy for Robotic Manipulation
- PointVLA: Injecting the 3D World into Vision-Language-Action Models
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
- DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping
- LLM Post-Training: A Deep Dive into Reasoning Large Language Models
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
- FetchBot: Learning Generalizable Object Fetching in Cluttered Scenes via Zero-Shot Sim2Real
- ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
- Magma: A Foundation Model for Multimodal AI Agents
- Pre-training Auto-regressive Robotic Models with 4D Representations
- AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors
- Scaling Pre-training to One Hundred Billion Data for Vision Language Models
- DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
- ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy
- Action-Free Reasoning for Policy Generalization
- From Foresight to Forethought: VLM-In-the-Loop Policy Steering via Latent Alignment
- UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
- Improving Vision-Language-Action Model with Online Reinforcement Learning
- SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- Universal Actions for Enhanced Embodied Foundation Models
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
- Towards Generalist Robot Policies: What Matters in Building Vision-Language-Action Models
- Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
- RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation
- Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
- Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos
- TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies
- RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning
- FLIP: Flow-Centric Generative Planning as General-Purpose Manipulation World Model
- RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation
- Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
- CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
- GRAPE: Generalizing Robot Policy via Preference Alignment
- Inference-Time Policy Steering through Human Interactions
- VQA2: Visual Question Answering for Video Quality Assessment
- DexMimicGen: Automated Data Generation for Bimanual Dexterous Manipulation via Imitation Learning
- π0: A Vision-Language-Action Flow Model for General Robot Control
- Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance
- Latent Action Pretraining from Videos
- RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
- Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation
- GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
- Generalizing Consistency Policy to Visual RL with Prioritized Proximal Experience Regularization
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation
- TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
- HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers
- All Robots in One: A New Standard and Unified Dataset for Versatile, General-Purpose Embodied Agents
- Robotic Control via Embodied Chain-of-Thought Reasoning
- PaliGemma: A versatile 3B VLM for transfer
- OpenVLA: An Open-Source Vision-Language-Action Model
- RVT-2: Learning Precise Manipulation from Few Demonstrations
- RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
- RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
- A Survey on Vision-Language-Action Models for Embodied AI
- Octo: An Open-Source Generalist Robot Policy
- Evaluating Real-World Robot Manipulation Policies in Simulation
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- 3D-VLA: A 3D Vision-Language-Action Generative World Model
- Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots
- FMB: a Functional Manipulation Benchmark for Generalizable Robotic Learning
- FMB: A functional manipulation benchmark for generalizable robotic learning
- SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers
- Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation
- RobotGPT: Robot Manipulation Learning from ChatGPT
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
- Vision-Language Foundation Models as Effective Robot Imitators
- CapsFusion: Rethinking Image-Text Data at Scale
- Boosting Continuous Control with Consistency Policy
- RoboHive: A Unified Framework for Robot Learning
- BridgeData V2: A Dataset for Robot Learning at Scale
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot
- RVT: Robotic View Transformer for 3D Object Manipulation
- LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
- DINOv2: Learning Robust Visual Features without Supervision
- Sigmoid Loss for Language Image Pre-Training
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
- Diffusion policy: Visuomotor policy learning via action diffusion
- PaLM-E: An Embodied Multimodal Language Model
- LLaMA: Open and Efficient Foundation Language Models
- ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
- Learning Universal Policies via Text-Guided Video Generation
- RT-1: Robotics Transformer for Real-World Control at Scale
- Behavior Transformers: Cloning k modes with one stone
- A Generalist Agent
- Flamingo: a Visual Language Model for Few-Shot Learning
- CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- Ego4D: Around the World in 3,000 Hours of Egocentric Video
- Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets
- CLIPort: What and Where Pathways for Robotic Manipulation
- What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
- Learning Transferable Visual Models From Natural Language Supervision
- Taming Transformers for High-Resolution Image Synthesis
- Transporter Networks: Rearranging the Visual World for Robotic Manipulation
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Language-Conditioned Imitation Learning for Robot Manipulation Tasks
- Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning
- Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
- RLBench: The Robot Learning Benchmark & Learning Environment
- A Short Note on the Kinetics-700 Human Action Dataset
- EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
- Towards VQA Models That Can Read
- Scaling Egocentric Vision: The EPIC-KITCHENS Dataset
- FiLM: Visual Reasoning with a General Conditioning Layer
- The "something something" video database for learning and evaluating visual common sense
- Attention Is All You Need
- Microsoft COCO: Common Objects in Context
- Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization
- A Taxonomy for Evaluating Generalist Robot Manipulation Policies
- Gemini Robotics: Bringing AI into the Physical World
- Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
Cited by
Related