Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
2025/11/03 by Xiaoyu Zhan, Zhan, Xiaoyu, Wenxuan Huang +25 · 3 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Constraint Satisfaction and Optimization #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2511.01618
openalex publication_date 2025/11/03 · openalex created_date 2025/11/06 · openalex updated_date 2026/07/28
Abstract
Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved 2D visual understanding, prompting interest in their application to complex 3D reasoning tasks. However, it remains unclear whether these models can effectively capture the detailed spatial information required for robust real-world performance, especially cross-view consistency, a key requirement for accurate 3D reasoning. Considering this issue, we introduce Viewpoint Learning, a task designed to evaluate and improve the spatial reasoning capabilities of MLLMs. We present the Viewpoint-100K dataset, consisting of 100K object-centric image pairs with diverse viewpoints and corresponding question-answer pairs. Our approach employs a two-stage fine-tuning strategy: first, foundational knowledge is injected to the baseline MLLM via Supervised Fine-Tuning (SFT) on Viewpoint-100K, resulting in significant improvements across multiple tasks; second, generalization is enhanced through Reinforcement Learning using the Group Relative Policy Optimization (GRPO) algorithm on a broader set of questions. Additionally, we introduce a hybrid cold-start initialization method designed to simultaneously learn viewpoint representations and maintain coherent reasoning thinking. Experimental results show that our approach significantly activates the spatial reasoning ability of MLLM, improving performance on both in-domain and out-of-domain reasoning tasks. Our findings highlight the value of developing foundational spatial skills in MLLMs, supporting future progress in robotics, autonomous systems, and 3D scene understanding.
Citations
- MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
- SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
- LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?
- MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation
- MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
- VGGT: Visual Geometry Grounded Transformer
- R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
- Qwen2.5-VL Technical Report
- ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning
- Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- 3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
- An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models
- GPT-4o System Card
- Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Multi-modal Situated Reasoning in 3D Scenes
- VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- SpatialBot: Precise Spatial Understanding with Vision Language Models
- Situational Awareness Matters in 3D Vision Language Reasoning
- BLINK: Multimodal Large Language Models Can See but Not Perceive
- Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
- Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
- Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- 3DAxiesPrompts: Unleashing the 3D Spatial Task Capabilities of GPT-4V
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- What's "up" with vision-language models? Investigating their struggle with spatial reasoning
- MVImgNet: A Large-scale Dataset of Multi-view Images
- Visual Spatial Reasoning
- Proximal Policy Optimization Algorithms
Cited by
Related