Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
2025/04/21 by Yeh, Chun-Hsiao, Wang, Chenyu, Tong, Shengbang +7
#Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2504.15280
Abstract
Multi-view understanding, the ability to reconcile visual information across diverse viewpoints for effective navigation, manipulation, and 3D scene comprehension, is a fundamental challenge in Multi-Modal Large Language Models (MLLMs) to be used as embodied agents. While recent MLLMs have shown impressive advances in high-level reasoning and planning, they frequently fall short when confronted with multi-view geometric consistency and cross-view correspondence. To comprehensively evaluate the challenges of MLLMs in multi-view scene reasoning, we propose All-Angles Bench, a benchmark of over 2,100 human carefully annotated multi-view question-answer pairs across 90 diverse real-world scenes. Our six tasks (counting, attribute identification, relative distance, relative direction, object manipulation, and camera pose estimation) specifically test model's geometric correspondence and the capacity to align information consistently across views. Our extensive experiments, benchmark on 27 representative MLLMs including Gemini-2.0-Flash, Claude-3.7-Sonnet, and GPT-4o against human evaluators reveals a substantial performance gap, indicating that current MLLMs remain far from human-level proficiency. Through in-depth analysis, we show that MLLMs are particularly underperforming under two aspects: (1) cross-view correspondence for partially occluded views and (2) establishing the coarse camera poses. These findings highlight the necessity of domain-specific refinements or modules that embed stronger multi-view awareness. We believe that our All-Angles Bench offers valuable insights and contribute to bridging the gap between MLLMs and human-level multi-view understanding. The project and benchmark are publicly available at https://danielchyeh.github.io/All-Angles-Bench/.
Citations
- Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning
- Forgotten Polygons: Multimodal Large Language Models are Shape-Blind
- Qwen2.5-VL Technical Report
- VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- Qwen2.5 Technical Report
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- HourVideo: 1-Hour Video-Language Understanding
- CAD-MLLM: Unifying Multimodality-Conditioned CAD Generation With MLLM
- GPT-4o System Card
- Dynamic Multimodal Evaluation with Flexible Complexity by Vision-Language Bootstrapping
- Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
- LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
- LongVILA: Scaling Long-Context Visual Language Models for Long Videos
- LLaVA-OneVision: Easy Visual Task Transfer
- VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
- DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning
- OpenVLA: An Open-Source Vision-Language-Action Model
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- Ovis: Structural Embedding Alignment for Multimodal Large Language Model
- Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
- Dynamic Evaluation of Large Language Models by Meta Probing Agents
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
- MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning
- SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
- MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- Improved Baselines with Visual Instruction Tuning
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
- MMBench: Is Your Multi-modal Model an All-around Player?
- EgoHumans: An Egocentric 3D Multi-Human Benchmark
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- Visual Instruction Tuning
- 3D Concept Learning and Reasoning from Multi-View Images
- ConceptFusion: Open-set Multimodal 3D Mapping
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents
- Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
- Multi-Target Embodied Question Answering
- Embodied Question Answering
Related