vix.ing · top · new · best · stats

InSpire: Vision-Language-Action Models with Intrinsic Spatial Reasoning

2025/05/20 by Ji Zhang, Zhang, Ji, Shihan Wu +11 · 10 citations
Computer Science · Social Sciences · #FOS: Computer and information sciences #Geographic Information Systems Studies #Multimodal Machine Learning Applications #Robotics (cs.RO) #Speech and dialogue systems

paper · pdf · doi:10.48550/arxiv.2505.13888

openalex publication_date 2025/05/20 · openalex created_date 2025/10/19 · openalex updated_date 2026/07/28

Abstract

Leveraging pretrained Vision-Language Models (VLMs) to map language instruction and visual observations to raw low-level actions, Vision-Language-Action models (VLAs) hold great promise for achieving general-purpose robotic systems. Despite their advancements, existing VLAs tend to spuriously correlate task-irrelevant visual features with actions, limiting their generalization capacity beyond the training data. To tackle this challenge, we propose Intrinsic Spatial Reasoning (InSpire), a simple yet effective approach that mitigates the adverse effects of spurious correlations by boosting the spatial reasoning ability of VLAs. Specifically, InSpire redirects the VLA's attention to task-relevant factors by prepending the question "In which direction is the [object] relative to the robot?" to the language instruction and aligning the answer "right/left/up/down/front/back/grasped" and predicted actions with ground-truth. Notably, InSpire can be used as a plugin to enhance existing autoregressive VLAs, requiring no extra training data or interaction with other large models. Extensive experimental results in both simulation and real-world environments demonstrate the effectiveness and flexibility of our approach.

Citations

Cited by

Related