VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
2025/04/10 by Haozhe Wang, Wang, Haozhe, Chao Qu +9 · 324 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2504.08837
openalex publication_date 2025/04/10 · openalex created_date 2025/10/14 · openalex updated_date 2026/07/28
Abstract
Recently, slow-thinking systems like GPT-o1 and DeepSeek-R1 have demonstrated great potential in solving challenging problems through explicit reflection. They significantly outperform the best fast-thinking models, such as GPT-4o, on various math and science benchmarks. However, their multimodal reasoning capabilities remain on par with fast-thinking models. For instance, GPT-o1's performance on benchmarks like MathVista, MathVerse, and MathVision is similar to fast-thinking models. In this paper, we aim to enhance the slow-thinking capabilities of vision-language models using reinforcement learning (without relying on distillation) to advance the state of the art. First, we adapt the GRPO algorithm with a novel technique called Selective Sample Replay (SSR) to address the vanishing advantages problem. While this approach yields strong performance, the resulting RL-trained models exhibit limited self-reflection or self-verification. To further encourage slow-thinking, we introduce Forced Rethinking, which appends a rethinking trigger token to the end of rollouts in RL training, explicitly enforcing a self-reflection reasoning step. By combining these two techniques, our model, VL-Rethinker, advances state-of-the-art scores on MathVista, MathVerse to achieve 80.4%, 63.5% respectively. VL-Rethinker also achieves open-source SoTA on multi-disciplinary benchmarks such as MathVision, MMMU-Pro, EMMA, and MEGA-Bench, narrowing the gap with OpenAI-o1. Our empirical results show the effectiveness of our approaches.
Citations
Cited by
- Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
- VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
- See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
- Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
- LanteRn: Latent Visual Structured Reasoning
- StAR: Segment Anything Reasoner
- CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
- Anatomy-R1: Enhancing Anatomy Reasoning in Multimodal Large Language Models via Anatomical Similarity Curriculum and Group Diversity Augmentation
- Stable and Efficient Single-Rollout RL for Multimodal Reasoning
- Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
- UniRel: Relation-Centric Knowledge Graph Question Answering with RL-Tuned LLM Reasoning
- AdaTooler-V: Adaptive Tool-Use for Images and Videos
- Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
- PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
- ADHint: Adaptive Hints with Difficulty Priors for Reinforcement Learning
- Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
- CogDoc: Towards Unified thinking in Documents
- Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
- Journey Before Destination: On the importance of Visual Faithfulness in Slow Thinking
- CADMorph: Geometry-Driven Parametric CAD Editing via a Plan-Generate-Verify Loop
- Enhancing Radiology Report Generation and Visual Grounding using Reinforcement Learning
- Rethinking Chain-of-Thought Reasoning for Videos
- ReLaX: Reasoning with Latent Exploration for Large Reasoning Models
- Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
- Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning
- OneThinker: All-in-one Reasoning Model for Image and Video
- VACoT: Rethinking Visual Data Augmentation with VLMs
- From Illusion to Intention: Visual Rationale Learning for Vision-Language Reasoning
- Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization
- Asking like Socrates: Socrates helps VLMs understand remote sensing images
- REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding
- SPHINX: A Synthetic Environment for Visual Perception and Reasoning
- Boosting Reasoning in Large Multimodal Models via Activation Replay
- VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
- VADE: Variance-Aware Dynamic Sampling via Online Sample-Level Difficulty Estimation for Multimodal RL
- Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
- Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT
- Learning to Think Fast and Slow for Visual Language Models
- OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
- Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling
- Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs
- Multimodal LLMs Do Not Compose Skills Optimally Across Modalities
- DeepEyesV2: Toward Agentic Multimodal Model
- SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
- Ariadne: A Controllable Framework for Probing and Extending VLM Reasoning Boundaries
- Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
- ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
- PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
- Rethinking the Text-Vision Reasoning Imbalance in MLLMs through the Lens of Training Recipes
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- Metis-HOME: Hybrid Optimized Mixture-of-Experts for Multimodal Reasoning
- Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
- AutoRubric-R1V: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
- Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
- HoneyBee: Data Recipes for Vision-Language Reasoners
- Scaling Language-Centric Omnimodal Representation Learning
- A Survey on Agentic Multimodal Large Language Models
- OmniQuality-R: Advancing Reward Models Through All-Encompassing Quality Assessment
- OpusAnimation: Code-Based Dynamic Chart Generation
- RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
- Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
- Spotlight on Token Perception for Multimodal Reinforcement Learning
- ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping
- SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models
- ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
- TTRV: Test-Time Reinforcement Learning for Vision Language Models
- What MLLMs Learn about When they Learn about Multimodal Reasoning
- Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images
- ACPO: Adaptive Curriculum Policy Optimization for Aligning Vision-Language Models in Complex Reasoning
- Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding
- Latent Visual Reasoning
- Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
- MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
- VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
- F2RVLM: Boosting Fine-grained Fragment Retrieval for Multi-Modal Long-form Dialogue with Vision Language Model
- DF-LLaVA: Unlocking MLLMs for Synthetic Image Detection via Knowledge Injection and Conflict-Driven Self-Reflection
- SAIL-VL2 Technical Report
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
- Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
- CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
- Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
- Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization
- Reverse-Engineered Reasoning for Open-Ended Generation
- Language-Driven Object-Oriented Two-Stage Method for Scene Graph Anticipation
- Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
- Attributes as Textual Genes: Leveraging LLMs as Genetic Algorithm Simulators for Conditional Synthetic Data Generation
- LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
- Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
- Bridging Formal Language with Chain-of-Thought Reasoning to Geometry Problem Solving
- MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models
- Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
- StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models
- AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
- VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
- EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
- Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner
- The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
- Perception-Aware Policy Optimization for Multimodal Reasoning
- Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
- Look-Back: Implicit Visual Re-focusing in MLLM Reasoning
- Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
- EFRame: Deeper Reasoning via Exploration-Filter-Replay Reinforcement Learning Framework
- BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
- DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO
- ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
- Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models
- WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
- Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?
- PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning
- Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
- Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
- VGR: Visual Grounded Reasoning
- Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation
- Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
- Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
- LUT: Latent Utility Training for Visual Reasoning
- Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning
- SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
- Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
- MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book
- MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning
- ProxyThinker: Test-Time Guidance through Small Visual Reasoners
- MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning
- X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
- InfiMed: Low-Resource Medical MLLMs with Advancing Understanding and Reasoning
- Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models
- Jigsaw-R1: A Study of Rule-based Visual Reinforcement Learning with Jigsaw Puzzles
- Sherlock: Self-Correcting Reasoning in Vision-Language Models
- R1-Code-Interpreter: LLMs Reason with Code via Supervised and Multi-stage Reinforcement Learning
- DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
- VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Gudied Iterative Policy Optimization
- SATORI-R1: Incentivizing Multimodal Reasoning through Explicit Visual Anchoring
- VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
- Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
- One RL to See Them All: Visual Triple Unified Reinforcement Learning
- Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL
- More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
- OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning
- R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
- Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
- UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
- Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning
- Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning
- Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models
- RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
- Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
- R3: Replay, Reflection, and Ranking Rewards for LLM Reinforcement Learning
- RelayLLM: Efficient Reasoning via Collaborative Decoding
- Aligning Large Vision-Language Models at Test Time: A Trajectory-Guided Structured Sampling Approach
- Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
Related