VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
2025/06/07 by Li, Can, Liu, Ying, Zhang, Ting +2 · 3 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2506.06727
Abstract
Large Multimodal Models have achieved remarkable progress in integrating vision and language, enabling strong performance across perception, reasoning, and domain-specific tasks. However, their capacity to reason over multiple, visually similar inputs remains insufficiently explored. Such fine-grained comparative reasoning is central to real-world tasks, especially in mathematics and education, where learners must often distinguish between nearly identical diagrams to identify correct solutions. To address this gap, we present VisioMath, a curated benchmark of 1,800 high-quality K-12 mathematics problems in which all candidate answers are diagrams with subtle visual similarities. A comprehensive evaluation of state-of-the-art LMMs, covering both leading closed-source systems and widely adopted open-source models, reveals a consistent decline in accuracy as inter-image similarity increases. Analysis indicates that the dominant failure mode stems from image-text misalignment: rather than grounding reasoning in textual cues, models often resort to shallow positional heuristics, resulting in systematic errors. We further explore three alignment-oriented strategies, spanning training-free approaches and finetuning, and achieve substantial accuracy gains. We hope that VisioMath will serve as a rigorous benchmark and catalyst for developing LMMs toward deeper diagram understanding, precise comparative reasoning, and grounded multi-image-text integration.
Citations
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision
- CPGD: Toward Stable Rule-based Reinforcement Learning for Language Models
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation
- R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
- MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- MV-MATH: Evaluating Multimodal Math Reasoning in Multi-Visual Contexts
- MMSciBench: Benchmarking Language Models on Chinese Multimodal Scientific Problems
- Qwen2.5-VL Technical Report
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models
- A Survey on Evaluation of Multimodal Large Language Models
- Building and better understanding vision-language models: insights and future directions
- Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning
- A Survey on Benchmarks of Multimodal Large Language Models
- SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning
- LLaVA-OneVision: Easy Visual Task Transfer
- The Llama 3 Herd of Models
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
- We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- From Pixels to Prose: A Large Dataset of Dense Image Captions
- MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
- OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification
- MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
- RelationVLM: Making Large Vision-Language Models Understand Visual Relations
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
- Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models
- CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
- Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models
- Evaluating the Performance of Large Language Models on GAOKAO Benchmark
- Visual Instruction Tuning
- Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text
- Training Verifiers to Solve Math Word Problems
- Measuring Mathematical Problem Solving With the MATH Dataset
- Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
- Unlocking Multimodal Mathematical Reasoning via Process Reward Model
Cited by
Related