SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
2024/01/22 by Boyuan Chen, Zhuo Xu, Chen, Boyuan +15 · 178 citations
Computer Science · Social Sciences · #Advanced Image and Video Retrieval Techniques #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Geographic Information Systems Studies #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Robotics (cs.RO)
paper · pdf · doi:10.48550/arxiv.2401.12168
openalex publication_date 2024/01/22 · openalex created_date 2024/01/24 · openalex updated_date 2026/07/28
Abstract
Understanding and reasoning about spatial relationships is a fundamental capability for Visual Question Answering (VQA) and robotics. While Vision Language Models (VLM) have demonstrated remarkable performance in certain VQA benchmarks, they still lack capabilities in 3D spatial reasoning, such as recognizing quantitative relationships of physical objects like distances or size differences. We hypothesize that VLMs' limited spatial reasoning capability is due to the lack of 3D spatial knowledge in training data and aim to solve this problem by training VLMs with Internet-scale spatial reasoning data. To this end, we present a system to facilitate this approach. We first develop an automatic 3D spatial VQA data generation framework that scales up to 2 billion VQA examples on 10 million real-world images. We then investigate various factors in the training recipe, including data quality, training pipeline, and VLM architecture. Our work features the first internet-scale 3D spatial reasoning dataset in metric space. By training a VLM on such data, we significantly enhance its ability on both qualitative and quantitative spatial VQA. Finally, we demonstrate that this VLM unlocks novel downstream applications in chain-of-thought spatial reasoning and robotics due to its quantitative estimation capability. Project website: https://spatial-vlm.github.io/
Cited by
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
- The Spatial Blindspot of Vision-Language Models
- GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding
- Teaching Tiny VLA Models Where to Look and How to Move
- Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning
- Semantic Deception: When Reasoning Models Can't Compute an Addition
- Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models
- From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
- Visually Prompted Benchmarks Are Surprisingly Fragile
- Neuro-Symbolic Control with Large Language Models for Language-Guided Spatial Tasks
- InfSplign: Inference-Time Spatial Alignment of Text-to-Image Diffusion Models
- N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
- Scaling Spatial Reasoning in MLLMs through Programmatic Data Synthesis
- EagleVision: A Dual-Stage Framework with BEV-grounding-based Chain-of-Thought for Spatial Intelligence
- Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
- Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos
- Lemon: A Unified and Scalable 3D Multimodal Model for Universal Spatial Understanding
- The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
- From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models
- MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
- SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
- Prompt-to-Parts: Generative AI for Physical Assembly and Scalable Instructions
- PAVAS: Physics-Aware Video-to-Audio Synthesis
- SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
- FRIEDA: Benchmarking Multi-Step Cartographic Reasoning in Vision-Language Models
- Towards Accurate UAV Image Perception: Guiding Vision-Language Models with Stronger Task Prompts
- DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
- Stitch and Tell: A Structured Multimodal Data Augmentation Method for Spatial Understanding
- SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models
- Towards Cross-View Point Correspondence in Vision-Language Models
- SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL
- Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
- MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm
- SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead
- Describe Anything Anywhere At Any Moment
- Obstruction reasoning for robotic grasping
- SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- Vision-Language Memory for Spatial Reasoning
- WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
- MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
- ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
- MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models
- Weakly-supervised Latent Models for Task-specific Visual-Language Control
- InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity
- RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
- IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
- SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding
- BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
- Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning
- First Frame Is the Place to Go for Video Content Customization
- C2F-Space: Coarse-to-Fine Space Grounding for Spatial Instructions using Vision-Language Models
- Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation
- Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation
- Video Spatial Reasoning with Object-Centric 3D Rollout
- RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation
- Imagine in Space: Exploring the Frontier of Spatial Intelligence and Reasoning Efficiency in Vision Language Models
- Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
- Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation
- Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric Refinement
- SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
- FractalBench: Diagnosing Visual-Mathematical Reasoning Through Recursive Program Synthesis
- iFlyBot-VLM Technical Report
- Visual Spatial Tuning
- Cambrian-S: Towards Spatial Supersensing in Video
- SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
- SVG Decomposition for Enhancing Large Multimodal Models Visualization Comprehension: A Study with Floor Plans
- An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM
- URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model
- OMEGA: Optimized Multimodal Position Encoding Index Derivation with Global Adaptive Scaling for Vision-Language Models
- Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
- A Multi-Modal Neuro-Symbolic Approach for Spatial Reasoning-Based Visual Grounding in Robotics
- Using VLM Reasoning to Constrain Task and Motion Planning
- Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
- Multimodal Language Models Cannot Spot Spatial Inconsistencies
- Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions
- Windsock is Dancing: Adaptive Multimodal Retrieval-Augmented Generation
- CityRiSE: Reasoning Urban Socio-Economic Status in Vision-Language Models via Reinforcement Learning
- GRAID: Enhancing Spatial Reasoning of VLMs Through High-Fidelity Data Generation
- Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation
- Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study
- PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
- Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
- Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
- HouseTour: A Virtual Real Estate A(I)gent
- GACO-CAD: Geometry-Augmented and Conciseness-Optimized CAD Model Generation from Single Image
- SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D Scenes
- Pursuing Minimal Sufficiency in Spatial Reasoning
- XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
- Talking Points: Describing and Localizing Pixels
- Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
- DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training
- Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph
- Prompt-Guided Spatial Understanding with RGB-D Transformers for Fine-Grained Object Relation Reasoning
- Holistic Order Prediction in Natural Scenes
- SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
- Unleashing Perception-Time Scaling to Multimodal Reasoning Models
- GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models
- TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
- EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
- Efficient Navigation in Unknown Indoor Environments with Vision-Language Models
- Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs
- GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert
- Spatial-ViLT: Enhancing Visual Spatial Reasoning through Multi-Task Learning
- OWL: Geometry-Aware Spatial Reasoning for Audio Large Language Models
- LLM-RG: Referential Grounding in Outdoor Scenarios using Large Language Models
- Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning Synergy
- DepthLM: Metric Depth From Vision Language Models
- Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
- RAVEN: Resilient Aerial Navigation via Open-Set Semantic Memory and Behavior Adaptation
- CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding
- FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion Editing
- See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation
- JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation
- RAU: Reference-based Anatomical Understanding with Vision Language Models
- CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
- TUN3D: Towards Real-World Scene Understanding from Unposed Images
- SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
- JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
- Gaze Heads: How VLMs Look at What They Describe
- Long Story Short: Disentangling Compositionality and Long-Caption Understanding in VLMs
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- The Case for Negative Data: From Crash Reports to Counterfactuals for Reasonable Driving
- SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
- MCTS-EP: Empowering Embodied Planning with Online Preference Optimization
- V-CECE: Visual Counterfactual Explanations via Conceptual Edits
- TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints
- A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
- Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation
- FiLM-Nav: Efficient and Generalizable Navigation via VLM Fine-tuning
- Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI
- SmolRGPT: Efficient Spatial Reasoning for Warehouse Environments with 600M Parameters
- CRAFT: Coaching Reinforcement Learning Autonomously using Foundation Models for Multi-Robot Coordination Tasks
- GestOS: Advanced Hand Gesture Interpretation via Large Language Models to control Any Type of Robot
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
- EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing
- 3D Aware Region Prompted Vision Language Model
- LVLMs are Bad at Overhearing Human Referential Communication
- Pathological Truth Bias in Vision-Language Models
- M3DMap: Object-aware Multimodal 3D Mapping for Dynamic Environments
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation
- Reconstruction Alignment Improves Unified Multimodal Models
- Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
- Towards Open World Detection: A Survey
- Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks
- DriveQA: Passing the Driving Knowledge Test
- How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images
- "Does the cafe entrance look accessible? Where is the door?" Towards Geospatial AI Agents for Visual Inquiries
- Spatial Policy: Guiding Visuomotor Robotic Manipulation with Spatial-Aware Modeling and Reasoning
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- See it. Say it. Sorted: Agentic System for Compositional Diagram Generation
- RynnEC: Bringing MLLMs into Embodied World
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
- CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance
- Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
- STRIDE-QA: Visual Question Answering Dataset for Spatiotemporal Reasoning in Urban Driving Scenes
- ORBIT: An Object Property Reasoning Benchmark for Visual Inference Tasks
- Spatial-ORMLLM: Improve Spatial Relation Understanding in the Operating Room with Multimodal Large Language Model
- AdjustAR: AI-Driven In-Situ Adjustment of Site-Specific Augmented Reality Content
- PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation
- SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- HOLODECK 2.0: Vision-Language-Guided 3D World Generation with Editing
- INTENTION: Inferring Tendencies of Humanoid Robot Motion Through Interactive Intuition and Grounded VLM
- NavA3: Understanding Any Instruction, Navigating Anywhere, Finding Anything
- Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models
- CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks
- From Pixels to Places: A Systematic Benchmark for Evaluating Image Geolocalization Ability in Large Language Models
- Context-based Motion Retrieval using Open Vocabulary Methods for Autonomous Driving
Related