SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
2025/10/09 by Li, Hongxing, Li, Dingming, Wang, Zixuan +7 · 8 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2510.08531
Abstract
Spatial reasoning remains a fundamental challenge for Vision-Language Models (VLMs), with current approaches struggling to achieve robust performance despite recent advances. We identify that this limitation stems from a critical gap: existing methods attempt to learn spatial reasoning directly without establishing the hierarchical foundations of perception and understanding. To address this challenge, we present a comprehensive methodology for building spatial intelligence progressively. We introduce SpatialLadder-26k, a multimodal dataset containing 26,610 samples spanning object localization, single image, multi-view, and video spatial reasoning tasks, constructed through a standardized pipeline that ensures systematic coverage across modalities. Building on this dataset, we design a three-stage progressive training framework that (1) establishes spatial perception through object localization, (2) develops spatial understanding through multi-dimensional spatial tasks, and (3) strengthens complex reasoning via reinforcement learning with verifiable rewards. This approach yields SpatialLadder, a 3B-parameter model that achieves state-of-the-art performance on spatial reasoning benchmarks, with 23.4% average improvement over the base model, surpassing GPT-4o by 20.8% and Gemini-2.0-Flash by 10.1%. Notably, SpatialLadder maintains strong generalization with 7.2% improvement on out-of-domain benchmarks, demonstrating that progressive training from perception to reasoning is essential for robust spatial intelligence.
Citations
- Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
- Self-Rewarding Vision-Language Model via Reasoning Decomposition
- Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
- Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
- ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
- GRIT: Teaching MLLMs to Think with Images
- Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
- Perception-R1: Pioneering Perception Policy with Reinforcement Learning
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
- Kimi-VL Technical Report
- SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
- Improved Visual-Spatial Reasoning via R1-Zero-Like Training
- STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
- From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D
- Video-R1: Reinforcing Video Reasoning in MLLMs
- VGGT: Visual Geometry Grounded Transformer
- R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- Visual-RFT: Visual Reinforcement Fine-Tuning
- Qwen2.5-VL Technical Report
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
- HourVideo: 1-Hour Video-Language Understanding
- GPT-4o System Card
- LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
- LLaVA-OneVision: Easy Visual Task Transfer
- Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- Mixtral of Experts
- What's "up" with vision-language models? Investigating their struggle with spatial reasoning
- ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild
- Training language models to follow instructions with human feedback
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data
- ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
- Microsoft COCO: Common Objects in Context
- SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
- 3D-LLM: Injecting the 3D World into Large Language Models
Cited by
Related