Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
2025/05/24 by Haoyuan Sun, Jiaqi Wu, Sun, Haoyuan +17 · 6 citations
Computer Science · #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2505.18536
Abstract
Standing in 2025, at a critical juncture in the pursuit of Artificial General Intelligence (AGI), reinforcement fine-tuning (RFT) has demonstrated significant potential in enhancing the reasoning capability of large language models (LLMs) and has led to the development of cutting-edge AI models such as OpenAI-o1 and DeepSeek-R1. Moreover, the efficient application of RFT to enhance the reasoning capability of multimodal large language models (MLLMs) has attracted widespread attention from the community. In this position paper, we argue that reinforcement fine-tuning powers the reasoning capability of multimodal large language models. To begin with, we provide a detailed introduction to the fundamental background knowledge that researchers interested in this field should be familiar with. Furthermore, we meticulously summarize the improvements of RFT in powering reasoning capability of MLLMs into five key points: diverse modalities, diverse tasks and domains, better training algorithms, abundant benchmarks and thriving engineering frameworks. Finally, we propose five promising directions for future research that the community might consider. We hope that this position paper will provide valuable insights to the community at this pivotal stage in the advancement toward AGI. Summary of works done on RFT for MLLMs is available at https://github.com/Sun-Haoyuan23/Awesome-RL-based-Reasoning-MLLMs.
Citations
- R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning
- OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
- DanceGRPO: Unleashing GRPO on Visual Generation
- Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning
- Flow-GRPO: Training Flow Matching Models via Online RL
- EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
- Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
- X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
- R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
- 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
- Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models
- GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling
- ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification
- Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
- Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
- SARI: Structured Audio Reasoning via Curriculum-Guided Reinforcement Learning
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
- Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark
- InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
- Compile Scene Graphs with Reinforcement Learning
- NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation
- GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning
- Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning
- SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
- GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
- TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning
- Perception-R1: Pioneering Perception Policy with Reinforcement Learning
- SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
- VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
- Kimi-VL Technical Report
- Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought
- MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models
- Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme
- Improved Visual-Spatial Reasoning via R1-Zero-Like Training
- Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1
- CrowdVLM-R1: Expanding R1 Ability to Vision Language Model for Crowd Counting using Fuzzy Group Relative Policy Reward
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
- Q-Insight: Understanding Image Quality via Visual Reinforcement Learning
- Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
- Video-R1: Reinforcing Video Reasoning in MLLMs
- Understanding R1-Zero-Like Training: A Critical Perspective
- MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
- Mind with Eyes: from Language Reasoning to Multimodal Reasoning
- OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning
- Think or Not Think: A Study of Explicit Thinking in Rule-Based Visual Reinforcement Fine-Tuning
- Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
- Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
- Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
- Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering
- VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
- Thinking Machines: A Survey of LLM based Reasoning Strategies
- Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
- R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
- Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
- Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning
- LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
- R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning
- R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
- Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
- What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
- Visual-RFT: Visual Reinforcement Fine-Tuning
- MV-MATH: Evaluating Multimodal Math Reasoning in Multi-Visual Contexts
- MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning
- A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos
- The Hidden Risks of Large Reasoning Models: A Safety Assessment of R1
- Investigating Inference-time Scaling for Chain of Multi-modal Thought: A Preliminary Study
- SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
- A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1
- MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
- ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
- Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
- iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
- MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models
- Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
- Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
- Virgo: A Preliminary Exploration on Reproducing o1-like MLLM
- Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
- Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
- Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
- DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
- Multi-Step Reasoning with Large Language Models, a Survey
- We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
- MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language Models
- Self-Supervised Visual Preference Alignment
- MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification
- LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models
- MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
- Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation
- Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- Safety of Multimodal Large Language Models on Images and Texts
- MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Aligning Large Multimodal Models with Factually Augmented RLHF
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- GPT-4 Technical Report
- LLaMA: Open and Efficient Foundation Language Models
- Training language models to follow instructions with human feedback
- Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning
- Rainbow: Combining Improvements in Deep Reinforcement Learning
- Proximal Policy Optimization Algorithms
- Concrete Problems in AI Safety
- Asynchronous Methods for Deep Reinforcement Learning
- Dueling Network Architectures for Deep Reinforcement Learning
- Deep Reinforcement Learning with Double Q-learning
- High-Dimensional Continuous Control Using Generalized Advantage\n Estimation
- Trust Region Policy Optimization
- Playing Atari with Deep Reinforcement Learning
- A Markovian Decision Process
- Relation-R1: Progressively Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relation Comprehension
- OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
- Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models
Cited by
Related