ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
2025/05/28 by Zhou, Zhongyi, Zhu, Yichen, Wen, Junjie +2 · 33 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics (cs.RO)
paper · doi:10.48550/arxiv.2505.21906
Abstract
Vision-language-action (VLA) models have emerged as the next generation of models in robotics. However, despite leveraging powerful pre-trained Vision-Language Models (VLMs), existing end-to-end VLA systems often lose key capabilities during fine-tuning as the model adapts to specific robotic tasks. We argue that a generalizable VLA model should retain and expand upon the VLM's core competencies: 1) Open-world embodied reasoning - the VLA should inherit the knowledge from VLM, i.e., recognize anything that the VLM can recognize, be capable of solving math problems, and possess visual-spatial intelligence, 2) Reasoning following - effectively translating the open-world reasoning into actionable steps for the robot. In this work, we introduce ChatVLA-2, a novel mixture-of-expert VLA model coupled with a specialized two-stage training pipeline designed to preserve the VLM's original strengths while enabling actionable reasoning. To validate our approach, we design a math-matching task wherein a robot interprets math problems written on a whiteboard and picks corresponding number cards from a table to solve equations. Remarkably, our method exhibits exceptional mathematical reasoning and OCR capabilities, despite these abilities not being explicitly trained within the VLA. Furthermore, we demonstrate that the VLA possesses strong spatial reasoning skills, enabling it to interpret novel directional instructions involving previously unseen objects. Overall, our method showcases reasoning and comprehension abilities that significantly surpass state-of-the-art imitation learning methods such as OpenVLA, DexVLA, and pi-zero. This work represents a substantial advancement toward developing truly generalizable robotic foundation models endowed with robust reasoning capacities.
Citations
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data
- OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
- π0.5: a Vision-Language-Action Model with Open-World Generalization
- Reasoning Models Can Be Effective Without Thinking
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
- PointVLA: Injecting the 3D World into Vision-Language-Action Models
- ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
- Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration
- ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
- DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control
- ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
- Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
- RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation
- Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
- CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
- Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation
- RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
- DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution
- π0: A Vision-Language-Action Flow Model for General Robot Control
- ALOHA Unleashed: A Simple Recipe for Robot Dexterity
- The Ingredients for Robotic Diffusion Transformers
- RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
- Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation
- Discrete Policy: Learning Disentangled Action Space for Multi-Task Robotic Manipulation
- Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic Manipulation
- TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Robotic Control via Embodied Chain-of-Thought Reasoning
- PaliGemma: A versatile 3B VLM for transfer
- Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
- Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning
- OpenVLA: An Open-Source Vision-Language-Action Model
- RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
- Octo: An Open-Source Generalist Robot Policy
- Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation
- Retrieval-Augmented Embodied Agents
- Feedback Efficient Online Fine-Tuning of Diffusion Models
- Fine-Tuning of Continuous-Time Diffusion Models as Entropy-Regularized Control
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
- Any-point Trajectory Modeling for Policy Learning
- QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
- An Embodied Generalist Agent in 3D World
- Vision-Language Foundation Models as Effective Robot Imitators
- Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Training Diffusion Models with Reinforcement Learning
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
- Diffusion policy: Visuomotor policy learning via action diffusion
- RT-1: Robotics Transformer for Real-World Control at Scale
- Visual Reinforcement Learning with Self-Supervised 3D Representations
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Towards VQA Models That Can Read
- Microsoft COCO: Common Objects in Context
- Gemini Robotics: Bringing AI into the Physical World
- Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
Cited by
Related