InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
2025/07/23 by S. Yang, Hao Li, Yang, Shuai +17 · 21 citations
Engineering · #Robotics and Automated Systems
paper · pdf · doi:10.48550/arxiv.2507.17520
Abstract
To operate effectively in the real world, robots should integrate multimodal reasoning with precise action generation. However, existing vision-language-action (VLA) models often sacrifice one for the other, narrow their abilities to task-specific manipulation data, and suffer catastrophic forgetting of pre-trained vision-language capabilities. To bridge this gap, we introduce InstructVLA, an end-to-end VLA model that preserves the flexible reasoning of large vision-language models (VLMs) while delivering leading manipulation performance with the help of embodied reasoning. InstructVLA introduces a novel training paradigm, Vision-Language-Action Instruction Tuning (VLA-IT), which employs multimodal training with mixture-of-experts adaptation to jointly optimize embodied reasoning and action generation on both standard VLM corpora and a curated 650K-sample VLA-IT dataset. On in-domain SimplerEnv tasks, InstructVLA achieves 33% improvement over SpatialVLA. To evaluate generalization, we introduce SimplerEnv-Instruct, an 80-task benchmark requiring closed-loop control and high-level instruction understanding, where it outperforms a fine-tuned OpenVLA by 96% and an action expert aided by GPT-4o by 29%. Additionally, InstructVLA surpasses baseline VLMs on multimodal tasks and exhibits inference-time scaling by leveraging textual reasoning to boost manipulation performance in both simulated and real-world settings. These results demonstrate InstructVLA's potential for bridging intuitive and steerable human-robot interaction with efficient policy learning.
Citations
- A Pragmatic VLA Foundation Model
- LongCat-Flash Technical Report
- GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
- Emerging Properties in Unified Multimodal Pretraining
- GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data
- RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
- π0.5: a Vision-Language-Action Model with Open-World Generalization
- Transfer between Modalities with MetaQueries
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- A Comprehensive Survey of Mixture-of-Experts: Algorithms, Theory, and Applications
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
- ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
- Magma: A Foundation Model for Multimodal AI Agents
- Pre-training Auto-regressive Robotic Models with 4D Representations
- HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation
- SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
- Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
- Universal Actions for Enhanced Embodied Foundation Models
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation
- Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
- CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
- π0: A Vision-Language-Action Flow Model for General Robot Control
- Latent Action Pretraining from Videos
- Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
- VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
- Robotic Control via Embodied Chain-of-Thought Reasoning
- PaliGemma: A versatile 3B VLM for transfer
- LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
- LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
- OpenVLA: An Open-Source Vision-Language-Action Model
- Octo: An Open-Source Generalist Robot Policy
- Evaluating Real-World Robot Manipulation Policies in Simulation
- From LLMs to Actions: Latent Codes as Bridges in Hierarchical Robot Control
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- RH20T-P: A Primitive-Level Robotic Dataset Towards Composable Generalization Agents
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- RT-H: Action Hierarchies Using Language
- Efficient Multimodal Learning from Data-centric Perspective
- Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
- X-LoRA: Mixture of Low-Rank Adapter Experts, a Flexible Framework for Large Language Models with Applications in Protein Mechanics and Molecular Design
- PoCo: Policy Composition from and for Heterogeneous Robot Learning
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches
- HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- MMBench: Is Your Multi-modal Model an All-around Player?
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
- Visual Instruction Tuning
- DINOv2: Learning Robust Visual Features without Supervision
- Sigmoid Loss for Language Image Pre-Training
- GPT-4 Technical Report
- LLaMA: Open and Efficient Foundation Language Models
- Scalable Diffusion Models with Transformers
- RT-1: Robotics Transformer for Real-World Control at Scale
- Flow Matching for Generative Modeling
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- Mixture-of-Experts with Expert Choice Routing
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Bridge Data: Boosting Generalization of Robotic Skills with Cross-Domain Datasets
- InfographicVQA
- SimCSE: Simple Contrastive Learning of Sentence Embeddings
- SimCSE: Simple Contrastive Learning of Sentence Embeddings
- Learning Transferable Visual Models From Natural Language Supervision
- DocVQA: A Dataset for VQA on Document Images
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- Towards VQA Models That Can Read
- FiLM: Visual Reasoning with a General Conditioning Layer
- A Diagram Is Worth A Dozen Images
- CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
- Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
- RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Cited by
Related