Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
2025/03/09 by Wenxuan Huang, Bohan Jia, Huang, Wenxuan +16 · 220 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2503.06749
openalex publication_date 2025/03/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
DeepSeek-R1-Zero has successfully demonstrated the emergence of reasoning capabilities in LLMs purely through Reinforcement Learning (RL). Inspired by this breakthrough, we explore how RL can be utilized to enhance the reasoning capability of MLLMs. However, direct training with RL struggles to activate complex reasoning capabilities such as questioning and reflection in MLLMs, due to the absence of substantial high-quality multimodal reasoning data. To address this issue, we propose the reasoning MLLM, Vision-R1, to improve multimodal reasoning capability. Specifically, we first construct a high-quality multimodal CoT dataset without human annotations by leveraging an existing MLLM and DeepSeek-R1 through modality bridging and data filtering to obtain a 200K multimodal CoT dataset, Vision-R1-cold dataset. It serves as cold-start initialization data for Vision-R1. To mitigate the optimization challenges caused by overthinking after cold start, we propose Progressive Thinking Suppression Training (PTST) strategy and employ Group Relative Policy Optimization (GRPO) with the hard formatting result reward function to gradually refine the model's ability to learn correct and complex reasoning processes on a 10K multimodal math dataset. Comprehensive experiments show our model achieves an average improvement of ∼6% across various multimodal math reasoning benchmarks. Vision-R1-7B achieves a 73.5% accuracy on the widely used MathVista benchmark, which is only 0.4% lower than the leading reasoning model, OpenAI O1. Scaling up the amount of multimodal math data in the RL training, Vision-R1-32B and Vison-R1-72B achieves 76.4% and 78.2% MathVista benchmark scores, respectively. The datasets and code will be released in: https://github.com/Osilly/Vision-R1 .
Cited by
- Offline-Online Curriculum RL for Multimodal Reasoning
- See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
- Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
- MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
- Latent Implicit Visual Reasoning
- CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
- ESearch-R1: Learning Cost-Aware MLLM Agents for Interactive Embodied Search via Reinforcement Learning
- Stable and Efficient Single-Rollout RL for Multimodal Reasoning
- UniRel: Relation-Centric Knowledge Graph Question Answering with RL-Tuned LLM Reasoning
- AdaTooler-V: Adaptive Tool-Use for Images and Videos
- Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
- PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
- A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
- ADHint: Adaptive Hints with Difficulty Priors for Reinforcement Learning
- CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
- DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model
- Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
- More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
- IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
- Enhancing Radiology Report Generation and Visual Grounding using Reinforcement Learning
- Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention
- Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation
- Rethinking Chain-of-Thought Reasoning for Videos
- Thinking with Images via Self-Calling Agent
- ReLaX: Reasoning with Latent Exploration for Large Reasoning Models
- START: Spatial and Textual Learning for Chart Understanding
- RVLF: A Reinforcing Vision-Language Framework for Gloss-Free Sign Language Translation
- Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning
- Knowing the Answer Isn't Enough: Fixing Reasoning Path Failures in LVLMs
- TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning
- OneThinker: All-in-one Reasoning Model for Image and Video
- GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes
- GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
- Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- Artemis: Structured Visual Reasoning for Perception Policy Learning
- OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
- JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization
- Asking like Socrates: Socrates helps VLMs understand remote sensing images
- Thinking in 360°: Humanoid Visual Search in the Wild
- Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
- LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
- Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
- Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
- SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization
- Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
- Learning to Think Fast and Slow for Visual Language Models
- Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval
- SkinGPT-R1: Adapter-Only Dual Distillation for Efficient Dermatology Reasoning
- ViSS-R1: Self-Supervised Reinforcement Video Reasoning
- From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
- MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation
- Simple Vision-Language Math Reasoning via Rendered Text
- Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
- From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
- SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
- Visual Spatial Tuning
- V-Thinker: Interactive Thinking with Images
- LaRe: Latent Refocusing for Multimodal Reasoning
- Ariadne: A Controllable Framework for Probing and Extending VLM Reasoning Boundaries
- VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
- PreferThinker: Reasoning-based Personalized Image Preference Assessment
- RoboOS-NeXT: A Unified Memory-based Framework for Lifelong, Scalable, and Robust Multi-Robot Collaboration
- ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
- Urban-R1: Reinforced MLLMs Mitigate Geospatial Biases for Urban General Intelligence
- Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
- ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
- VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
- Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
- FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
- Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
- UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoning
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- VAR: Visual Attention Reasoning via Structured Search and Backtracking
- Proactive Reasoning-with-Retrieval Framework for Medical Multimodal Large Language Models
- Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- Scaling Language-Centric Omnimodal Representation Learning
- CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
- A Survey on Agentic Multimodal Large Language Models
- OmniQuality-R: Advancing Reward Models Through All-Encompassing Quality Assessment
- ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models
- MARS-Sep: Multimodal-Aligned Reinforced Sound Separation
- Answer-Consistent Chain-of-thought Reinforcement Learning For Multi-modal Large Langauge Models
- RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
- Spotlight on Token Perception for Multimodal Reinforcement Learning
- Towards Efficient Multimodal Unified Reasoning Model via Model Merging
- Unleashing Perception-Time Scaling to Multimodal Reasoning Models
- SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
- Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
- SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models
- What MLLMs Learn about When they Learn about Multimodal Reasoning
- λ-GRPO: Unifying the GRPO Frameworks with Learnable Token Preferences
- HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection
- Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images
- Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
- Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models
- Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces
- BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
- TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Language Models
- VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding
- Latent Visual Reasoning
- Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO
- GeoVLM-R1: Reinforcement Fine-Tuning for Improved Remote Sensing Reasoning
- ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis
- Advancing Multi-agent Traffic Simulation via R1-Style Reinforcement Fine-Tuning
- EditGRPO: Reinforcement Learning with Post-Rollout Edits for Clinically Accurate Chest X-Ray Report Generation
- MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
- ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models
- RISK: A Framework for GUI Agents in E-commerce Risk Management
- Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning
- Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
- Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
- Learning GUI Grounding with Spatial Reasoning from Visual Feedback
- MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
- GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
- Orcust: Stepwise-Feedback Reinforcement Learning for GUI Agent
- Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language Models
- Advancing Speech Understanding in Speech-Aware Language Models with GRPO
- SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
- ChartMaster: Advancing Chart-to-Code Generation with Real-World Charts and Chart Similarity Reinforcement Learning
- Generalizable Geometric Image Caption Synthesis
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- Tool-R1: Sample-Efficient Reinforcement Learning for Agentic Tool Use
- Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
- Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
- Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
- Reconstruction Alignment Improves Unified Multimodal Models
- Teaching AI Stepwise Diagnostic Reasoning with Report-Guided Chain-of-Thought Learning
- Interleaving Reasoning for Better Text-to-Image Generation
- Reinforced Visual Perception with Tools
- Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal Reasoning
- LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
- InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
- Self-Rewarding Vision-Language Model via Reasoning Decomposition
- Do MLLMs Really Understand the Charts?
- Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
- Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking in Large Speech Models
- Thyme: Think Beyond Images
- UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning
- We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
- MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
- Bridging Formal Language with Chain-of-Thought Reasoning to Geometry Problem Solving
- Time Is a Feature: Exploiting Temporal Dynamics in Diffusion Language Models
- Audio-Thinker: Guiding Audio Language Model When and How to Think via Reinforcement Learning
- Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
- Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments
- AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
- Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
- A Survey of Multimodal Ophthalmic Diagnostics: From Task-Specific Approaches to Foundational Models
- VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
- FairReason: Balancing Reasoning and Social Bias in MLLMs
- Position: Reasoning After Perception Means Reasoning Without Vision
- OctoNav: Towards Generalist Embodied Navigation
- Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
- Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner
- DiagR1: A Vision-Language Model Trained via Reinforcement Learning for Digestive Pathology Diagnosis
- Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
- Benchmarking Gaslighting Negation Attacks Against Reasoning Models
- PositionIC: Unified Position and Identity Consistency for Image Customization
- Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning
- C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
- ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
- VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
- M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
- Scaling RL to Long Videos
- Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
- The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
- Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
- Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs
- Perception-Aware Policy Optimization for Multimodal Reasoning
- High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
- R1-RE: Cross-Domain Relation Extraction with RLVR
- Multimodal Mathematical Reasoning with Diverse Solving Perspective
- Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs
- Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
- VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
- Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency
- Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
- MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
- Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling
- SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents
- APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
- HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
- DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO
- HiMA-Ecom: Enabling Joint Training of Hierarchical Multi-Agent E-commerce Assistants
- GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior
- RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models
- Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning
- WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
- JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent
- Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
- GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
- PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning
- Play to Generalize: Learning to Reason Through Game Play
- Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
- Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
- Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
- VGR: Visual Grounded Reasoning
- Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning
- Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation
- Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
- Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward
Related