MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
2023/10/03 by Lu, Pan, Bansal, Hritik, Xia, Tony +7 · 358 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
paper · doi:10.48550/arxiv.2310.02255
Abstract
Large Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive problem-solving skills in many tasks and domains, but their ability in mathematical reasoning in visual contexts has not been systematically studied. To bridge this gap, we present MathVista, a benchmark designed to combine challenges from diverse mathematical and visual tasks. It consists of 6,141 examples, derived from 28 existing multimodal datasets involving mathematics and 3 newly created datasets (i.e., IQTest, FunctionQA, and PaperQA). Completing these tasks requires fine-grained, deep visual understanding and compositional reasoning, which all state-of-the-art foundation models find challenging. With MathVista, we have conducted a comprehensive, quantitative evaluation of 12 prominent foundation models. The best-performing GPT-4V model achieves an overall accuracy of 49.9%, substantially outperforming Bard, the second-best performer, by 15.1%. Our in-depth analysis reveals that the superiority of GPT-4V is mainly attributed to its enhanced visual perception and mathematical reasoning. However, GPT-4V still falls short of human performance by 10.4%, as it often struggles to understand complex figures and perform rigorous reasoning. This significant gap underscores the critical role that MathVista will play in the development of general-purpose AI agents capable of tackling mathematically intensive and visually rich real-world tasks. We further explore the new ability of self-verification, the application of self-consistency, and the interactive chatbot capabilities of GPT-4V, highlighting its promising potential for future research. The project is available at https://mathvista.github.io/.
Cited by
- Same or Not? Enhancing Visual Perception in Vision-Language Models
- VL-RouterBench: A Benchmark for Vision-Language Model Routing
- HiSciBench: A Hierarchical Multi-disciplinary Benchmark for Scientific Intelligence from Reading to Discovery
- GOTS: Greedy Orthogonal Token Selection for High-Resolution Vision-Language Models
- Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- Offline-Online Curriculum RL for Multimodal Reasoning
- Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation
- PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions
- See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
- StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design
- CHaystack: Benchmarking Chinese Document Retrieval and VQA
- When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
- UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
- Your Reasoning Benchmark May Not Test Reasoning: Revealing Perception Bottleneck in Abstract Reasoning Benchmarks
- Concept Generalization in Humans and Large Language Models: Insights from the Number Game
- CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
- OpenView: Empowering MLLMs with Out-of-view VQA
- Stable and Efficient Single-Rollout RL for Multimodal Reasoning
- FPBench: A Comprehensive Benchmark of Multimodal Large Language Models for Fingerprint Analysis
- Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding
- CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
- Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
- Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
- UmniBench: Unified Understand and Generation Model Oriented Omni-dimensional Benchmark
- AdaTooler-V: Adaptive Tool-Use for Images and Videos
- MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning
- Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
- The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
- DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
- Step-GUI Technical Report
- PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
- V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions
- ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
- JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
- HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
- ADHint: Adaptive Hints with Difficulty Priors for Reinforcement Learning
- STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
- CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
- FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
- Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
- More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
- Journey Before Destination: On the importance of Visual Faithfulness in Slow Thinking
- Limits and Gains of Test-Time Scaling in Vision-Language Reasoning
- Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention
- Towards Fine-Grained Recognition with Large Visual Language Models: Benchmark and Optimization Strategies
- CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
- PyFi: Toward Pyramid-like Financial Image Understanding for VLMs via Adversarial Agents
- Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules
- DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
- ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
- Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs
- InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
- FRIEDA: Benchmarking Multi-Step Cartographic Reasoning in Vision-Language Models
- ReLaX: Reasoning with Latent Exploration for Large Reasoning Models
- START: Spatial and Textual Learning for Chart Understanding
- Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
- Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- Knowing the Answer Isn't Enough: Fixing Reasoning Path Failures in LVLMs
- TRACE: A Framework for Analyzing and Enhancing Stepwise Reasoning in Vision-Language Models
- EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
- Qwen3.5-Omni Technical Report
- Jina-VLM: Small Multilingual Vision Language Model
- AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition
- OneThinker: All-in-one Reasoning Model for Image and Video
- Hierarchical Process Reward Models are Symbolic Vision Learners
- Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
- MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm
- PPTBench: Towards Holistic Evaluation of Large Language Models for PowerPoint Layout and Design Understanding
- Reframing Human-Robot Interaction Through Extended Reality: Unlocking Safer, Smarter, and More Empathic Interactions with Virtual Robots and Foundation Models
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering
- MathSight: A Benchmark Exploring Have Vision-Language Models Really Seen in University-Level Mathematical Reasoning?
- From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
- AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
- Structured Extraction from Business Process Diagrams Using Vision-Language Models
- Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
- Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning
- WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios
- Qwen3-VL Technical Report
- SPHINX: A Synthetic Environment for Visual Perception and Reasoning
- AlignBench: Benchmarking Fine-Grained Image-Text Alignment with Synthetic Image-Caption Pairs
- Boosting Reasoning in Large Multimodal Models via Activation Replay
- Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
- VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
- Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
- CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization
- Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
- OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
- DuoTeach: Dual Role Self-Teaching for Coarse-to-Fine Decision Coordination in Vision--Language Models
- Understanding Task Transfer in Vision-Language Models
- VADE: Variance-Aware Dynamic Sampling via Online Sample-Level Difficulty Estimation for Multimodal RL
- SO-Bench: A Structural Output Evaluation of Multimodal LLMs
- Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
- AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
- SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization
- RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
- L2V-CoT: Cross-Modal Transfer of Chain-of-Thought Reasoning via Latent Intervention
- Attention Guided Alignment in Efficient Vision-Language Models
- ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
- EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
- Learning to Think Fast and Slow for Visual Language Models
- When to Think and When to Look: Uncertainty-Guided Lookback
- Multimodal Evaluation of Russian-language Architectures
- CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution
- Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
- Stealth Fine-Tuning: Efficiently Breaking Alignment in RVLMs Using Self-Generated CoT
- TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
- From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- EcoAlign: An Economically Rational Framework for Efficient LVLM Alignment
- GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
- Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis for Large Reasoning Models
- Simple Vision-Language Math Reasoning via Rendered Text
- mmJEE-Eval: A Bilingual Multimodal Benchmark for Evaluating Scientific Reasoning in Vision-Language Models
- FaithAct: Faithfulness Planning and Acting in MLLMs
- DiagramIR: An Automatic Pipeline for Educational Math Diagram Evaluation
- MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
- Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish View
- FractalBench: Diagnosing Visual-Mathematical Reasoning Through Recursive Program Synthesis
- Unveiling Modality Bias: Automated Sample-Specific Analysis for Multimodal Misinformation Benchmarks
- DeepEyesV2: Toward Agentic Multimodal Model
- Cambrian-S: Towards Spatial Supersensing in Video
- Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
- V-Thinker: Interactive Thinking with Images
- NVIDIA Nemotron Nano V2 VL
- Contamination Detection for VLMs using Multi-Modal Semantic Perturbation
- MME-CC: A Challenging Multi-Modal Evaluation Benchmark of Cognitive Capacity
- ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
- Can Visual Input Be Compressed? A Visual Token Compression Benchmark for Large Multimodal Models
- ChartM3: A Multi-Stage Code-Driven Pipeline for Constructing Multi-Dimensional and Multi-Step Visual Reasoning Data in Chart Comprehension
- SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
- TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
- Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
- ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use
- LongCat-Flash-Omni Technical Report
- Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
- GeoFM: Enhancing Geometric Reasoning of MLLMs via Synthetic Data Generation through Formal Language
- MM-OPERA: Benchmarking Open-ended Association Reasoning for Large Vision-Language Models
- PairUni: Pairwise Training for Unified Multimodal Language Models
- Understanding Knowledge Transfer Mechanism in Heterogeneous MLLM Fusion: A Simple Linear Approach
- PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning
- MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
- RL makes MLLMs see better than SFT
- PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
- Ministral 3
- Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
- ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games?
- Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
- Latent Chain-of-Thought for Visual Reasoning
- PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
- VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
- MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding
- MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring
- PISA-Bench: The PISA Index as a Multilingual and Multimodal Metric for the Evaluation of Vision-Language Models
- DynaSolidGeo: A Dynamic Benchmark for Genuine Spatial Mathematical Reasoning of VLMs in Solid Geometry
- Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- Designing and Evaluating Hint Generation Systems for Science Education
- Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
- Diagnosing Visual Reasoning: Challenges, Insights, and a Path Forward
- GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
- Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
- Metis-HOME: Hybrid Optimized Mixture-of-Experts for Multimodal Reasoning
- [De|Re]constructing VLMs' Reasoning in Counting
- CARES: Context-Aware Resolution Selector for VLMs
- KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
- When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
- UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
- Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
- VERITAS: Leveraging Vision Priors and Expert Fusion to Improve Multimodal Data
- MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
- Directional Reasoning Injection for Fine-Tuning MLLMs
- Attention Is All You Need for KV Cache in Diffusion LLMs
- MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
- AutoRubric-R1V: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
- Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
- MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
- Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
- RECODE: Reasoning Through Code Generation for Visual Question Answering
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
- DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search
- MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science
- MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
- HoneyBee: Data Recipes for Vision-Language Reasoners
- ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
- GeoVLMath: Enhancing Geometry Reasoning in Vision-Language Models via Cross-Modal Reward for Auxiliary Line Creation
- A Survey on Agentic Multimodal Large Language Models
- CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
- VOGUE: Guiding Exploration with Visual Uncertainty Improves Multimodal Reasoning
- ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models
- MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
- RLFR: Extending Reinforcement Learning for LLMs with Flow Environment
- Answer-Consistent Chain-of-thought Reinforcement Learning For Multi-modal Large Langauge Models
- Reallocating Attention Across Layers to Reduce Multimodal Hallucination
- Towards Efficient Multimodal Unified Reasoning Model via Model Merging
- CapGeo: A Caption-Assisted Approach to Geometric Reasoning
- BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
- To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
- ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping
- Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception
- FinMR: A Knowledge-Intensive Multimodal Benchmark for Advanced Financial Reasoning
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks
- ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
- TTRV: Test-Time Reinforcement Learning for Vision Language Models
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- What MLLMs Learn about When they Learn about Multimodal Reasoning
- ARISE: An Adaptive Resolution-Aware Metric for Test-Time Scaling Evaluation in Large Reasoning Models
- Seeing the Big Picture: Evaluating Multimodal LLMs' Ability to Interpret and Grade Handwritten Student Work
- Beyond Monolithic Rewards: A Hybrid and Multi-Aspect Reward Optimization for MLLM Alignment
- COSMO-RL: Towards Trustworthy LMRMs via Joint Safety and Stability
- Self-Improvement in Multimodal Large Language Models: A Survey
- Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models
- Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-Level Physics Problem Solving
- ACPO: Adaptive Curriculum Policy Optimization for Aligning Vision-Language Models in Complex Reasoning
- MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles
- Apriel-1.5-15b-Thinker
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces
- Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
- DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
- MuSLR: Multimodal Symbolic Logical Reasoning
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
- Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking
- LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
- RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMs
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
- AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy
- Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
- LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models
- RIV: Recursive Introspection Mask Diffusion Vision Language Model
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
- WirelessMathLM: Teaching Mathematical Reasoning for LLMs in Wireless Communications with Reinforcement Learning
- MMPB: It's Time for Multi-Modal Personalization
- REMA: A Unified Reasoning Manifold Framework for Interpreting Large Language Model
- GeoSketch: A Neural-Symbolic Approach to Geometric Multimodal Reasoning with Auxiliary Line Construction and Affine Transformation
- Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding
- CircuitSense: A Hierarchical Circuit System Benchmark Bridging Visual Comprehension and Symbolic Reasoning in Engineering Design Process
- Multilingual Vision-Language Models, A Survey
- GenesisGeo: Technical Report
- Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
- Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
- CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
- MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
- Human-like Navigation in a World Built for Humans
- GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions
- FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification
- Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- Proximal Supervised Fine-Tuning
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
- MAPO: Mixed Advantage Policy Optimization
- Steering Multimodal Large Language Models Decoding for Context-Aware Safety
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving
- RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
- GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning
- Qwen3-Omni Technical Report
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
- BaseReward: A Strong Baseline for Multimodal Reward Model
- Generalizable Geometric Image Caption Synthesis
- AToken: A Unified Tokenizer for Vision
- Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
- Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
- The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations
- AsyMoE: Leveraging Modal Asymmetry for Enhanced Expert Specialization in Large Vision-Language Models
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
- Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
- 3D Aware Region Prompted Vision Language Model
- MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
- Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
- Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration
- Can Vision-Language Models Solve Visual Math Equations?
- Bringing Multi-Modal Multi-Task Federated Foundation Models to Education Domain: Prospects and Challenges
- DreamPRM-1.5: Unlocking the Potential of Each Instance for Multimodal Process Reward Model Training
- The Telephone Game: Evaluating Semantic Drift in Unified Models
- Promptception: How Sensitive Are Large Multimodal Models to Prompts?
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
- Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
- Reinforced Visual Perception with Tools
- Improving Large Vision and Language Models by Learning from a Panel of Peers
- Kwai Keye-VL 1.5 Technical Report
- Robix: A Unified Model for Robot Interaction, Reasoning and Planning
- LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
- The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?
- R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning
- Intern-S1: A Scientific Multimodal Foundation Model
- 11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
- PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark
- WP-CLIP: Leveraging CLIP to Predict Wölfflin's Principles in Visual Art
- Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
- Thyme: Think Beyond Images
- Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps
- Ovis2.5 Technical Report
- MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
- We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
- MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
- MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
- Bridging Formal Language with Chain-of-Thought Reasoning to Geometry Problem Solving
- Reinforcement Learning for Large Model: A Survey
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- Effective Training Data Synthesis for Improving MLLM Chart Understanding
- MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models
- GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines
- Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
- Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
- StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging
- GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning
- Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models
- Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability
- VRPRM: Process Reward Modeling via Visual Reasoning
- Geoint-R1: Formalizing Multimodal Geometric Reasoning with Dynamic Auxiliary Constructions
- A Rolling Stone Gathers No Moss: Adaptive Policy Optimization for Stable Self-Evaluation in Large Multimodal Models
- AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
- CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
- To Theoretically Understand Transformer-Based In-Context Learning for Optimizing CSMA
- GanitBench: A bi-lingual benchmark for evaluating mathematical reasoning in Vision Language Models
- Evaluating Contrast Localizer for Identifying Causal Units in Social & Mathematical Tasks in Language Models
- MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
- A Large Language Model Powered Integrated Circuit Footprint Geometry Understanding
- VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
- MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning
- CHECK-MAT: Checking Hand-Written Mathematical Answers for the Russian Unified State Exam
- Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
- SESR-Eval: Dataset for Evaluating LLMs in the Title-Abstract Screening of Systematic Reviews
- Wide-In, Narrow-Out: Revokable Decoding for Efficient and Effective DLLMs
- EH-Benchmark Ophthalmic Hallucination Benchmark and Agent-Driven Top-Down Traceable Reasoning Workflow
- MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
Related