SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
2025/05/22 by Haoning Wu, Wu, Haoning, Xiao Huang +8 · 12 citations
Computer Science · Engineering · #Multimodal Machine Learning Applications #Spatial Cognition and Navigation #Constraint Satisfaction and Optimization
paper · pdf · doi:10.48550/arxiv.2505.17012
Abstract
Existing evaluations of multimodal large language models (MLLMs) on spatial intelligence are typically fragmented and limited in scope. In this work, we aim to conduct a holistic assessment of the spatial understanding capabilities of modern MLLMs and propose complementary data-driven and agent-based solutions. Specifically, we make the following contributions: (i) we introduce SpatialScore, to our knowledge, the most comprehensive and diverse benchmark for multimodal spatial intelligence to date. It covers multiple visual data types, input modalities, and question-answering formats, and contains approximately 5K manually verified samples spanning 30 distinct tasks; (ii) using SpatialScore, we extensively evaluate 49 representative MLLMs, revealing persistent challenges and a substantial gap between current models and human-level spatial intelligence; (iii) to advance model capabilities, we construct SpatialCorpus, a large-scale training resource with 331K multimodal QA samples that supports fine-tuning on spatial reasoning tasks and significantly improves the performance of existing models (e.g., Qwen3-VL); (iv) to complement this data-driven route with a training-free paradigm, we develop SpatialAgent, a multi-agent system equipped with 12 specialized spatial perception tools that supports both Plan-Execute and ReAct reasoning, enabling substantial gains in spatial reasoning without additional model training. Extensive experiments and in-depth analyses demonstrate the effectiveness of our benchmark, corpus, and agent framework. We expect these resources to serve as a solid foundation for advancing MLLMs toward human-level spatial intelligence. All data, code, and models will be released to the research community.
Citations
- Qwen3-VL Technical Report
- Visual Spatial Tuning
- Cambrian-S: Towards Spatial Supersensing in Video
- SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
- Detect Anything via Next Point Prediction
- MapAnything: Universal Feed-Forward Metric 3D Reconstruction
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass
- gpt-oss-120b & gpt-oss-20b Model Card
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
- SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization
- MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
- SATORI-R1: Incentivizing Multimodal Reasoning through Explicit Visual Anchoring
- Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
- MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
- SITE: towards Spatial Intelligence Thorough Evaluation
- Multi-Agent System for Comprehensive Soccer Understanding
- SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
- SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Detect Anything 3D in the Wild
- Kimi-VL Technical Report
- SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
- STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
- From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D
- MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
- VGGT: Visual Geometry Grounded Transformer
- Qwen2.5-VL Technical Report
- Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
- PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
- Cosmos World Foundation Model Platform for Physical AI
- Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- 3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Cubify Anything: Scaling Indoor 3D Object Detection
- Towards Universal Soccer Video Understanding
- LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant
- RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
- LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
- DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
- SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models
- Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models
- Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks
- LLaVA-OneVision: Easy Visual Task Transfer
- MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
- SAM 2: Segment Anything in Images and Videos
- The Llama 3 Herd of Models
- MatchTime: Towards Automatic Soccer Game Commentary Generation
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models
- SpatialBot: Precise Spatial Understanding with Vision Language Models
- Depth Anything V2
- SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents
- BLINK: Multimodal Large Language Models Can See but Not Perceive
- MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
- Embodied LLM Agents Learn to Cooperate in Organized Teams
- Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
- RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
- DUSt3R: Geometric 3D Vision Made Easy
- VILA: On Pre-training for Visual Language Models
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- What's "up" with vision-language models? Investigating their struggle with spatial reasoning
- Improved Baselines with Visual Instruction Tuning
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes
- AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point Tracking
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- ChatDev: Communicative Agents for Software Development
- MMBench: Is Your Multi-modal Model an All-around Player?
- Visual Instruction Tuning
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild
- Visual Spatial Reasoning
- RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
- RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
- SpatialSense: An Adversarially Crowdsourced Benchmark for Spatial Relation Recognition
- ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
- Distinctive Image Features from Scale-Invariant Keypoints
- Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
- MindCube: Spatial Mental Modeling from Limited Views
- SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
- LATTE: Learning to Think with Vision Specialists
Cited by
Related