MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
2023/10/03 by Pan Lu, Lu, Pan, Hritik Bansal +17 · 653 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2310.02255
openalex publication_date 2023/10/03 · openalex created_date 2023/10/05 · openalex updated_date 2026/07/28
Abstract
Large Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive problem-solving skills in many tasks and domains, but their ability in mathematical reasoning in visual contexts has not been systematically studied. To bridge this gap, we present MathVista, a benchmark designed to combine challenges from diverse mathematical and visual tasks. It consists of 6,141 examples, derived from 28 existing multimodal datasets involving mathematics and 3 newly created datasets (i.e., IQTest, FunctionQA, and PaperQA). Completing these tasks requires fine-grained, deep visual understanding and compositional reasoning, which all state-of-the-art foundation models find challenging. With MathVista, we have conducted a comprehensive, quantitative evaluation of 12 prominent foundation models. The best-performing GPT-4V model achieves an overall accuracy of 49.9%, substantially outperforming Bard, the second-best performer, by 15.1%. Our in-depth analysis reveals that the superiority of GPT-4V is mainly attributed to its enhanced visual perception and mathematical reasoning. However, GPT-4V still falls short of human performance by 10.4%, as it often struggles to understand complex figures and perform rigorous reasoning. This significant gap underscores the critical role that MathVista will play in the development of general-purpose AI agents capable of tackling mathematically intensive and visually rich real-world tasks. We further explore the new ability of self-verification, the application of self-consistency, and the interactive chatbot capabilities of GPT-4V, highlighting its promising potential for future research. The project is available at https://mathvista.github.io/.
Cited by
- Same or Not? Enhancing Visual Perception in Vision-Language Models
- VL-RouterBench: A Benchmark for Vision-Language Model Routing
- HiSciBench: A Hierarchical Multi-disciplinary Benchmark for Scientific Intelligence from Reading to Discovery
- GOTS: Greedy Orthogonal Token Selection for High-Resolution Vision-Language Models
- Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- Offline-Online Curriculum RL for Multimodal Reasoning
- Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation
- PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions
- See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
- StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design
- CHaystack: Benchmarking Chinese Document Retrieval and VQA
- When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
- UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
- Your Reasoning Benchmark May Not Test Reasoning: Revealing Perception Bottleneck in Abstract Reasoning Benchmarks
- Concept Generalization in Humans and Large Language Models: Insights from the Number Game
- CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
- OpenView: Empowering MLLMs with Out-of-view VQA
- Stable and Efficient Single-Rollout RL for Multimodal Reasoning
- FPBench: A Comprehensive Benchmark of Multimodal Large Language Models for Fingerprint Analysis
- Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding
- CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
- Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
- Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
- UmniBench: Unified Understand and Generation Model Oriented Omni-dimensional Benchmark
- AdaTooler-V: Adaptive Tool-Use for Images and Videos
- MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning
- Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
- The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
- DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
- Step-GUI Technical Report
- PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
- V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions
- ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
- JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
- HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
- ADHint: Adaptive Hints with Difficulty Priors for Reinforcement Learning
- STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
- CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
- FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
- Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
- More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
- Journey Before Destination: On the importance of Visual Faithfulness in Slow Thinking
- Limits and Gains of Test-Time Scaling in Vision-Language Reasoning
- Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention
- Towards Fine-Grained Recognition with Large Visual Language Models: Benchmark and Optimization Strategies
- CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
- PyFi: Toward Pyramid-like Financial Image Understanding for VLMs via Adversarial Agents
- Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules
- DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
- ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
- Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs
- InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
- FRIEDA: Benchmarking Multi-Step Cartographic Reasoning in Vision-Language Models
- ReLaX: Reasoning with Latent Exploration for Large Reasoning Models
- START: Spatial and Textual Learning for Chart Understanding
- Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
- Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- Knowing the Answer Isn't Enough: Fixing Reasoning Path Failures in LVLMs
- TRACE: A Framework for Analyzing and Enhancing Stepwise Reasoning in Vision-Language Models
- EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
- Qwen3.5-Omni Technical Report
- Jina-VLM: Small Multilingual Vision Language Model
- AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition
- OneThinker: All-in-one Reasoning Model for Image and Video
- Hierarchical Process Reward Models are Symbolic Vision Learners
- Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
- MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm
- PPTBench: Towards Holistic Evaluation of Large Language Models for PowerPoint Layout and Design Understanding
- Reframing Human-Robot Interaction Through Extended Reality: Unlocking Safer, Smarter, and More Empathic Interactions with Virtual Robots and Foundation Models
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering
- MathSight: A Benchmark Exploring Have Vision-Language Models Really Seen in University-Level Mathematical Reasoning?
- From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
- AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
- Structured Extraction from Business Process Diagrams Using Vision-Language Models
- Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
- Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning
- WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios
- Qwen3-VL Technical Report
- SPHINX: A Synthetic Environment for Visual Perception and Reasoning
- HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
- Boosting Reasoning in Large Multimodal Models via Activation Replay
- Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
- VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
- Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
- CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization
- Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
- OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
- DuoTeach: Dual Role Self-Teaching for Coarse-to-Fine Decision Coordination in Vision--Language Models
- Understanding Task Transfer in Vision-Language Models
- VADE: Variance-Aware Dynamic Sampling via Online Sample-Level Difficulty Estimation for Multimodal RL
- SO-Bench: A Structural Output Evaluation of Multimodal LLMs
- Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
- AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
- SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization
- RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
- L2V-CoT: Cross-Modal Transfer of Chain-of-Thought Reasoning via Latent Intervention
- Attention Guided Alignment in Efficient Vision-Language Models
- ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
- EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
- Learning to Think Fast and Slow for Visual Language Models
- When to Think and When to Look: Uncertainty-Guided Lookback
- Multimodal Evaluation of Russian-language Architectures
- CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution
- Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
- Stealth Fine-Tuning: Efficiently Breaking Alignment in RVLMs Using Self-Generated CoT
- TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
- From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- EcoAlign: An Economically Rational Framework for Efficient LVLM Alignment
- GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
- Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis for Large Reasoning Models
- Simple Vision-Language Math Reasoning via Rendered Text
- mmJEE-Eval: A Bilingual Multimodal Benchmark for Evaluating Scientific Reasoning in Vision-Language Models
- Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs
- DiagramIR: An Automatic Pipeline for Educational Math Diagram Evaluation
- MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
- Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish View
- FractalBench: Diagnosing Visual-Mathematical Reasoning Through Recursive Program Synthesis
- Unveiling Modality Bias: Automated Sample-Specific Analysis for Multimodal Misinformation Benchmarks
- DeepEyesV2: Toward Agentic Multimodal Model
- Cambrian-S: Towards Spatial Supersensing in Video
- Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
- V-Thinker: Interactive Thinking with Images
- NVIDIA Nemotron Nano V2 VL
- Contamination Detection for VLMs using Multi-Modal Semantic Perturbation
- MME-CC: A Challenging Multi-Modal Evaluation Benchmark of Cognitive Capacity
- ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
- Can Visual Input Be Compressed? A Visual Token Compression Benchmark for Large Multimodal Models
- ChartM3: A Multi-Stage Code-Driven Pipeline for Constructing Multi-Dimensional and Multi-Step Visual Reasoning Data in Chart Comprehension
- SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
- TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
- Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
- ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use
- LongCat-Flash-Omni Technical Report
- Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
- GeoFM: Enhancing Geometric Reasoning of MLLMs via Synthetic Data Generation through Formal Language
- MM-OPERA: Benchmarking Open-ended Association Reasoning for Large Vision-Language Models
- PairUni: Pairwise Training for Unified Multimodal Language Models
- Understanding Knowledge Transfer Mechanism in Heterogeneous MLLM Fusion: A Simple Linear Approach
- PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning
- MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
- RL makes MLLMs see better than SFT
- PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
- Ministral 3
- Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
- ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games?
- Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
- Latent Chain-of-Thought for Visual Reasoning
- PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
- VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
- MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding
- MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring
- PISA-Bench: The PISA Index as a Multilingual and Multimodal Metric for the Evaluation of Vision-Language Models
- DynaSolidGeo: A Dynamic Benchmark for Genuine Spatial Mathematical Reasoning of VLMs in Solid Geometry
- Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- Designing and Evaluating Hint Generation Systems for Science Education
- Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
- Diagnosing Visual Reasoning: Challenges, Insights, and a Path Forward
- GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
- Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
- Metis-HOME: Hybrid Optimized Mixture-of-Experts for Multimodal Reasoning
- [De|Re]constructing VLMs' Reasoning in Counting
- CARES: Context-Aware Resolution Selector for VLMs
- KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
- When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
- UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
- Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
- VERITAS: Leveraging Vision Priors and Expert Fusion to Improve Multimodal Data
- MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
- Directional Reasoning Injection for Fine-Tuning MLLMs
- Attention Is All You Need for KV Cache in Diffusion LLMs
- MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
- AutoRubric-R1V: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
- Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
- MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
- Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
- RECODE: Reasoning Through Code Generation for Visual Question Answering
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
- DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search
- MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science
- MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
- HoneyBee: Data Recipes for Vision-Language Reasoners
- ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
- GeoVLMath: Enhancing Geometry Reasoning in Vision-Language Models via Cross-Modal Reward for Auxiliary Line Creation
- A Survey on Agentic Multimodal Large Language Models
- CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
- Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning
- ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models
- MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
- RLFR: Extending Reinforcement Learning for LLMs with Flow Environment
- Answer-Consistent Chain-of-thought Reinforcement Learning For Multi-modal Large Langauge Models
- Reallocating Attention Across Layers to Reduce Multimodal Hallucination
- Towards Efficient Multimodal Unified Reasoning Model via Model Merging
- CapGeo: A Caption-Assisted Approach to Geometric Reasoning
- BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
- To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
- ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping
- Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception
- FinMR: A Knowledge-Intensive Multimodal Benchmark for Advanced Financial Reasoning
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks
- ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
- TTRV: Test-Time Reinforcement Learning for Vision Language Models
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- What MLLMs Learn about When they Learn about Multimodal Reasoning
- ARISE: An Adaptive Resolution-Aware Metric for Test-Time Scaling Evaluation in Large Reasoning Models
- Seeing the Big Picture: Evaluating Multimodal LLMs' Ability to Interpret and Grade Handwritten Student Work
- Beyond Monolithic Rewards: A Hybrid and Multi-Aspect Reward Optimization for MLLM Alignment
- COSMO-RL: Towards Trustworthy LMRMs via Joint Safety and Stability
- Self-Improvement in Multimodal Large Language Models: A Survey
- Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models
- Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-Level Physics Problem Solving
- ACPO: Adaptive Curriculum Policy Optimization for Aligning Vision-Language Models in Complex Reasoning
- MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles
- Apriel-1.5-15b-Thinker
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces
- Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
- DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
- MuSLR: Multimodal Symbolic Logical Reasoning
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
- Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking
- LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
- RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMs
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
- AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy
- Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
- LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models
- RIV: Recursive Introspection Mask Diffusion Vision Language Model
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
- WirelessMathLM: Teaching Mathematical Reasoning for LLMs in Wireless Communications with Reinforcement Learning
- MMPB: It's Time for Multi-Modal Personalization
- REMA: A Unified Reasoning Manifold Framework for Interpreting Large Language Model
- GeoSketch: A Neural-Symbolic Approach to Geometric Multimodal Reasoning with Auxiliary Line Construction and Affine Transformation
- Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding
- CircuitSense: A Hierarchical Circuit System Benchmark Bridging Visual Comprehension and Symbolic Reasoning in Engineering Design Process
- Multilingual Vision-Language Models, A Survey
- GenesisGeo: Technical Report
- Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
- Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
- CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
- MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
- Human-like Navigation in a World Built for Humans
- GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions
- FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification
- Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- Proximal Supervised Fine-Tuning
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
- MAPO: Mixed Advantage Policy Optimization
- Steering Multimodal Large Language Models Decoding for Context-Aware Safety
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving
- RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
- GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning
- Qwen3-Omni Technical Report
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
- BaseReward: A Strong Baseline for Multimodal Reward Model
- Generalizable Geometric Image Caption Synthesis
- AToken: A Unified Tokenizer for Vision
- Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
- Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
- The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations
- AsyMoE: Leveraging Modal Asymmetry for Enhanced Expert Specialization in Large Vision-Language Models
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
- Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
- 3D Aware Region Prompted Vision Language Model
- MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
- Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
- Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration
- Can Vision-Language Models Solve Visual Math Equations?
- Bringing Multi-Modal Multi-Task Federated Foundation Models to Education Domain: Prospects and Challenges
- DreamPRM-1.5: Unlocking the Potential of Each Instance for Multimodal Process Reward Model Training
- The Telephone Game: Evaluating Semantic Drift in Unified Models
- Promptception: How Sensitive Are Large Multimodal Models to Prompts?
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
- Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
- Reinforced Visual Perception with Tools
- Improving Large Vision and Language Models by Learning from a Panel of Peers
- Kwai Keye-VL 1.5 Technical Report
- Robix: A Unified Model for Robot Interaction, Reasoning and Planning
- LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
- The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?
- R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning
- Intern-S1: A Scientific Multimodal Foundation Model
- 11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
- PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark
- WP-CLIP: Leveraging CLIP to Predict Wölfflin's Principles in Visual Art
- Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
- Thyme: Think Beyond Images
- Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps
- Ovis2.5 Technical Report
- MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
- We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
- MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
- MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
- Bridging Formal Language with Chain-of-Thought Reasoning to Geometry Problem Solving
- Reinforcement Learning for Large Model: A Survey
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- Effective Training Data Synthesis for Improving MLLM Chart Understanding
- MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models
- GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines
- Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
- Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
- StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models
- Quantifying uncert-AI-nty: Testing the accuracy of LLMs’ confidence judgments
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging
- GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning
- Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models
- Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability
- VRPRM: Process Reward Modeling via Visual Reasoning
- Geoint-R1: Formalizing Multimodal Geometric Reasoning with Dynamic Auxiliary Constructions
- A Rolling Stone Gathers No Moss: Adaptive Policy Optimization for Stable Self-Evaluation in Large Multimodal Models
- AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
- VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
- CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
- To Theoretically Understand Transformer-Based In-Context Learning for Optimizing CSMA
- GanitBench: A bi-lingual benchmark for evaluating mathematical reasoning in Vision Language Models
- Evaluating Contrast Localizer for Identifying Causal Units in Social & Mathematical Tasks in Language Models
- MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
- A Large Language Model Powered Integrated Circuit Footprint Geometry Understanding
- VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
- MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning
- CHECK-MAT: Checking Hand-Written Mathematical Answers for the Russian Unified State Exam
- Position: Reasoning After Perception Means Reasoning Without Vision
- Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
- SESR-Eval: Dataset for Evaluating LLMs in the Title-Abstract Screening of Systematic Reviews
- Wide-In, Narrow-Out: Revokable Decoding for Efficient and Effective DLLMs
- EH-Benchmark Ophthalmic Hallucination Benchmark and Agent-Driven Top-Down Traceable Reasoning Workflow
- MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
- The Invisible Leash: Why RLVR May or May Not Escape Its Origin
- MathDuels: Evaluating LLMs as Problem Posers and Solvers
- ERNIE 5.0 Technical Report
- Benchmarking Gaslighting Negation Attacks Against Reasoning Models
- Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning
- Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
- Pixels to Principles: Probing Intuitive Physics Understanding in Multimodal Language Models
- C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
- INTEGRALBENCH: Benchmarking LLMs with Definite Integral Problems
- A Survey of Deep Learning for Geometry Problem Solving
- Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
- Hyperphantasia: A Benchmark for Evaluating the Mental Visualization Capabilities of Multimodal LLMs
- ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
- VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism
- Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
- Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions
- M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
- Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought
- Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
- PyVision: Agentic Vision with Dynamic Tooling
- Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
- The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
- Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs
- Robust Multimodal Large Language Models Against Modality Conflict
- Perception-Aware Policy Optimization for Multimodal Reasoning
- Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling
- Skywork-R1V3 Technical Report
- BlueLM-2.5-3B Technical Report
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
- Vision-Language Models Can't See the Obvious
- MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding
- Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
- Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
- BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
- Multimodal Mathematical Reasoning with Diverse Solving Perspective
- Cautious Next Token Prediction
- Kwai Keye-VL Technical Report
- Look-Back: Implicit Visual Re-focusing in MLLM Reasoning
- SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
- From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought
- Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency
- MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
- Ovis-U1 Technical Report
- MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning
- EFRame: Deeper Reasoning via Exploration-Filter-Replay Reinforcement Learning Framework
- MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
- HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
- OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
- Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
- MMSearch-R1: Incentivizing LMMs to Search
- Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models
- Solving Inequality Proofs with Large Language Models
- CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning
- WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
- PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
- Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?
- AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
- Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
- Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning
- Demystifying the Visual Quality Paradox in Multimodal Large Language Models
- GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
- PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning
- Play to Generalize: Learning to Reason Through Game Play
- Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language Models
- SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
- MoTE: Mixture of Ternary Experts for Memory-efficient Large Multimodal Models
- Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
- Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
- GeoSDF: Plane Geometry Diagram Synthesis via Signed Distance Field
- GeometryZero: Improving Geometry Solving for LLM with Group Contrastive Policy Optimization
- Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
- FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design
- MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
- Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models
- MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval
- Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs
- Benchmarking Multimodal LLMs on Recognition and Understanding over Chemical Tables
- Magistral
- VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
- Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
- MMMG: A Massive, Multidisciplinary, Multi-Tier Generation Benchmark for Text-to-Image Reasoning
- Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?
- Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning
- Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
- Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
- Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
- MATP-BENCH: Can MLLM Be a Good Automated Theorem Prover for Multimodal Problems?
- FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging
- PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts
- CoMemo: LVLMs Need Image Context with Image Memory
- Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
- MLLM-CL: Continual Learning for Multimodal Large Language Models
- MMRefine: Unveiling the Obstacles to Robust Refinement in Multimodal Large Language Models
- When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
- Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Study
- Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
- VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
- MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
- More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
- MANBench: Is Your Multimodal Model Smarter than Human?
- Generating Pedagogically Meaningful Visuals for Math Word Problems: A New Benchmark and Analysis of Text-to-Image Models
- MiMo-VL Technical Report
- Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning
- PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models
- Towards Geometry Problem Solving in the Large Model Era: A Survey
- VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning
- SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis
- SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
- VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
- K12Vista: Exploring the Boundaries of MLLMs in K-12 Education
- NavBench: Probing Multimodal Large Language Models for Embodied Navigation
- Improve MLLM Benchmark Efficiency through Interview
- GuessBench: Sensemaking Multimodal Creativity in the Wild
- GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
- MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book
- EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models
- Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts
- CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs
- AMSbench: A Comprehensive Benchmark for Evaluating MLLM Capabilities in AMS Circuits
- ProxyThinker: Test-Time Guidance through Small Visual Reasoners
- MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning
- MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
- Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
- When Large Multimodal Models Confront Evolving Knowledge: Challenges and Explorations
- VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
- X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
- VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
- Elicit and Enhance: Advancing Multimodal Reasoning in Medical Scenarios
- Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models
- AutoGPS: Automated Geometry Problem Solving via Multimodal Formalization and Deductive Reasoning
- ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering
- Sherlock: Self-Correcting Reasoning in Vision-Language Models
- Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
- Self-Reflective Reinforcement Learning for Diffusion-based Image Reasoning Generation
- Do You See Me : A Multidimensional Benchmark for Evaluating Visual Perception in Multimodal LLMs
- NegVQA: Can Vision Language Models Understand Negation?
- MMTBENCH: A Unified Benchmark for Complex Multimodal Table Reasoning
- AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
- Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
- VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
- MMGeoLM: Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models
- Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models
- Can Visual Encoder Learn to See Arrows?
- Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution
- DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
- AI4Math: A Native Spanish Benchmark for University-Level Mathematical Reasoning in Large Language Models
- SATORI-R1: Incentivizing Multimodal Reasoning through Explicit Visual Anchoring
- DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative Decoding
- SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning
- Recompute or Reuse? Diagnosing and Mitigating Textual Shortcuts in VLM Self-Reflection
- v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
- Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
- Caption This, Reason That: VLMs Caught in the Middle
- Seeing Beyond Words: MatVQA for Challenging Visual-Scientific Reasoning in Materials Science
- One RL to See Them All: Visual Triple Unified Reinforcement Learning
- PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs
- Towards General Continuous Memory for Vision-Language Models
- GeoGramBench: Benchmarking the Geometric Program Reasoning in Modern LLMs
- Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
- More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
- CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
- Let Androids Dream of Electric Sheep: A Human-Inspired Image Implication Understanding and Reasoning Framework
- SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward
- OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning
- RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs
- Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning
- xInv: Explainable Optimization of Inverse Problems
- Dimple: Discrete Diffusion Multimodal Large Language Model with Parallel Decoding
- LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
- R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
- Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models
- Training-Free Reasoning and Reflection in MLLMs
- SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
- NAN: A Training-Free Solution to Coefficient Estimation in Model Merging
- Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
- ChartCards: A Chart-Metadata Generation Framework for Multi-Task Chart Understanding
- Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval
- GRIT: Teaching MLLMs to Think with Images
- Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems
- Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
- TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning
- Decouple and Orthogonalize: A Data-Free Framework for LoRA Merging
- Distill What the Student Can See: Fisher-Projected On-Policy Distillation for Vision-Language Models
- Emerging Properties in Unified Multimodal Pretraining
- Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey
- Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- Debating for Better Reasoning: An Unsupervised Multimodal Approach
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
- Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning
- Unlocking the Potential of Difficulty Prior in RL-based Multimodal Reasoning
- Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
- Beyond Routing Saturation: A Long-Horizon Class-Incremental Perspective on Expert Routing in Multimodal Continual Instruction Tuning
- MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision
- ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
- Advancing Sequential Numerical Prediction in Autoregressive Models
- Synthetic History: Evaluating Visual Representations of the Past in Diffusion Models
- LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
- Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning
- IQBench: How "Smart'' Are Vision-Language Models? A Study with Human IQ Tests
- Are LLMs Ready for English Standardized Tests? A Benchmarking and Elicitation Perspective
- LongChart VQA: A Comprehensive Benchmark for MLLMs with Complex Multi-Chart Reasoning
- Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
- MTRE: Multi-Token Reliability Estimation for Hallucination Detection in VLMs
- MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning
- Benchmarking LLMs on File System Design and Implementation
- Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput
- OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
- CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts
- What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs
- Critique Before Thinking: Mitigating Hallucination through Rationale-Augmented Instruction Tuning
- Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning
- Visual Instruction Tuning with Chain of Region-of-Interest
- CellVerse: Do Large Language Models Really Understand Cell Biology?
- Arrow-Guided VLM: Enhancing Flowchart Understanding via Arrow Direction Encoding
- Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design Understanding
- Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images
- R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
- R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning
- Kwai Keye-VL-2.0 Technical Report
- Infinity-Parser2 Technical Report
- IntrAgent: An LLM Agent for Content-Grounded Information Retrieval through Literature Review
- MiniMax Sparse Attention
- The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors
- ODAR: Principled Adaptive Routing for LLM Reasoning via Active Inference
- Can Generalist Agents Automate Data Curation?
- Beyond Language Modeling: An Exploration of Multimodal Pretraining
- HLL: Can Agents Cross Humanity's Last Line of Verification?
- ETCHR: Editing To Clarify and Harness Reasoning
- Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
- SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
- Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
- SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
- Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation
- OpenCoF: Learning to Reason Through Video Generation
- Large language model-enabled automated data extraction for concrete materials informatics
- TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training
- Let ViT Speak: Generative Language-Image Pre-training
- MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
- LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model
- RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
- PlotChain: Deterministic Checkpointed Evaluation of Multimodal LLMs on Engineering Plot Reading
- Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
- Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
- Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
- VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
- Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
- MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining
- BabyVision: Visual Reasoning Beyond Language
- When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation
- Does Forgetting Transfer Across Modalities? A Real-World Benchmark for Cross-Modal Knowledge Unlearning Evaluation
- ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
- Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency
- Representing Visual Evidence for Item Difficulty Prediction: Visual Textualization and Image-Native Modeling
- When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning
- Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
- GEB-Bench: Abstract Structures Told in Many Voices
- Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
- Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward
- TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
- Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
- Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark
- UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation
- Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?
- CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
- Benchmarking Vision Language Models on German Factual Data
- The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
- COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
- TAMP: Token-Adaptive Layerwise Pruning in Multimodal Large Language Models
- A Survey on Efficient Vision-Language Models
- PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Models
- VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
- SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
- Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models
- Kimi-VL Technical Report
- OmniCaptioner: One Captioner to Rule Them All
- SCI-Reason: A Dataset with Chain-of-Thought Rationales for Complex Multimodal Reasoning in Academic Areas
- MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models
- Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought
- SmolVLM: Redefining small and efficient multimodal models
Related