SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
2025/11/26 by Xu, Peiran, Wang, Sudong, Zhu, Yao +2 · 1 citation
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2511.21471
Abstract
Spatial cognition is fundamental to real-world multimodal intelligence, allowing models to effectively interact with the physical environment. While multimodal large language models (MLLMs) have made significant strides, existing benchmarks often oversimplify spatial cognition, reducing it to a single-dimensional metric, which fails to capture the hierarchical structure and interdependence of spatial abilities. To address this gap, we propose a hierarchical spatial cognition framework that decomposes spatial intelligence into five progressively complex levels from basic observation to high-level planning. Building upon this taxonomy, we construct SpatialBench, a large-scale, fine-grained benchmark covering 15 tasks aligned with these cognitive levels. To provide a unified evaluation across heterogeneous tasks, we further introduce a high-level capability-oriented metric that reliably assesses a model's overall spatial reasoning ability. Extensive experiments over massive MLLMs reveal distinct performance stratification across cognitive levels: models exhibit strong perceptual grounding yet remain limited in symbolic reasoning, causal inference, and planning. Additional human tests demonstrate that humans perform selective, goal-directed abstraction, while MLLMs tend to over-attend to surface details without coherent spatial intent. Our work establishes the first systematic framework for measuring hierarchical spatial cognition in MLLMs, laying the foundation for future spatially intelligent systems.
Citations
- Visual Spatial Tuning
- Cambrian-S: Towards Spatial Supersensing in Video
- SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- MindCube: Spatial Mental Modeling from Limited Views
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
- ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
- STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
- From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D
- ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos
- MiniMax-01: Scaling Foundation Models with Lightning Attention
- GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
- GPT-4o System Card
- Does Spatial Cognition Emerge in Frontier Models?
- LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
- Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
- LLaVA-OneVision: Easy Visual Task Transfer
- VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- SpatialBot: Precise Spatial Understanding with Vision Language Models
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
- SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
- Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
- GROUNDHOG: Grounding Large Language Models to Holistic Segmentation
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- Osprey: Pixel Understanding with Visual Instruction Tuning
- Honeybee: Locality-enhanced Projector for Multimodal LLM
- LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and Planning
- ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
- Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
- mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
- Evaluating Spatial Understanding of Large Language Models
- Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
- Improved Baselines with Visual Instruction Tuning
- InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
- DreamLLM: Synergistic Multimodal Comprehension and Creation
- EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
- GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
- LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
- Training Diffusion Models with Reinforcement Learning
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Visual Instruction Tuning
- ChatGPT outperforms crowd workers for text-annotation tasks
- GPT-4 Technical Report
- Diversity-Aware Meta Visual Prompting
- PaLM-E: An Embodied Multimodal Language Model
- LLaMA: Open and Efficient Foundation Language Models
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- InstructPix2Pix: Learning to Follow Image Editing Instructions
- SQA3D: Situated Question Answering in 3D Scenes
- PaLI: A Jointly-Scaled Multilingual Language-Image Model
- Flamingo: a Visual Language Model for Few-Shot Learning
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- ScanQA: 3D Question Answering for Spatial Scene Understanding
- Learning Transferable Visual Models From Natural Language Supervision
- Measuring Massive Multitask Language Understanding
- Language Models are Few-Shot Learners
- ERNIE: Enhanced Language Representation with Informative Entities
- Building Machines that Learn and Think for Themselves: Commentary on Lake et al., Behavioral and Brain Sciences, 2017
- Microsoft COCO: Common Objects in Context
- Vacunación del niño inmigrante y adoptado en España
- Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Cited by
Related