SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
2025/04/10 by Chen, Hardy, Haoqin Tu, Tu, Haoqin +12 · 94 citations
Computer Science · #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2504.11468
openalex publication_date 2025/04/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
This work revisits the dominant supervised fine-tuning (SFT) then reinforcement learning (RL) paradigm for training Large Vision-Language Models (LVLMs), and reveals a key finding: SFT can significantly undermine subsequent RL by inducing ``pseudo reasoning paths'' imitated from expert models. While these paths may resemble the native reasoning paths of RL models, they often involve prolonged, hesitant, less informative steps, and incorrect reasoning. To systematically study this effect, we introduce VLAA-Thinking, a new multimodal dataset designed to support reasoning in LVLMs. Constructed via a six-step pipeline involving captioning, reasoning distillation, answer rewrite and verification, VLAA-Thinking comprises high-quality, step-by-step visual reasoning traces for SFT, along with a more challenging RL split from the same data source. Using this dataset, we conduct extensive experiments comparing SFT, RL and their combinations. Results show that while SFT helps models learn reasoning formats, it often locks aligned models into imitative, rigid reasoning modes that impede further learning. In contrast, building on the Group Relative Policy Optimization (GRPO) with a novel mixed reward module integrating both perception and cognition signals, our RL approach fosters more genuine, adaptive reasoning behavior. Notably, our model VLAA-Thinker, based on Qwen2.5VL 3B, achieves top-1 performance on Open LMM Reasoning Leaderboard (https://huggingface.co/spaces/opencompass/OpenLMMReasoningLeaderboard) among 4B scale LVLMs, surpassing the previous state-of-the-art by 1.8%. We hope our findings provide valuable insights in developing reasoning-capable LVLMs and can inform future research in this area.
Cited by
- ProGuard: Towards Proactive Multimodal Safeguard
- MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning
- Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
- Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding
- Generalization of RLVR Using Causal Reasoning as a Testbed
- Stable and Efficient Single-Rollout RL for Multimodal Reasoning
- Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
- Trust-Region Adaptive Policy Optimization
- PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
- ADHint: Adaptive Hints with Difficulty Priors for Reinforcement Learning
- Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
- Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
- More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
- All You Need Are Random Visual Tokens? Demystifying Token Pruning in VLLMs
- Asking like Socrates: Socrates helps VLMs understand remote sensing images
- Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following
- The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation
- Boosting Reasoning in Large Multimodal Models via Activation Replay
- Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
- Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding
- Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
- MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models
- TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots
- OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
- Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
- SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
- Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
- DeepEyesV2: Toward Agentic Multimodal Model
- V-Thinker: Interactive Thinking with Images
- SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
- Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
- Urban-R1: Reinforced MLLMs Mitigate Geospatial Biases for Urban General Intelligence
- Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
- VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
- Think before Recommendation: Autonomous Reasoning-enhanced Recommender
- GeoThought: A Dataset for Enhancing Mathematical Geometry Reasoning in Vision-Language Models
- Metis-HOME: Hybrid Optimized Mixture-of-Experts for Multimodal Reasoning
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
- A Survey on Agentic Multimodal Large Language Models
- Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales
- Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
- Unleashing Perception-Time Scaling to Multimodal Reasoning Models
- ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping
- Customer-R1: Personalized Simulation of Human Behaviors via RL-based LLM Agent in Online Shopping
- Presenting a Paper is an Art: Self-Improvement Aesthetic Agents for Academic Presentations
- Beyond Monolithic Rewards: A Hybrid and Multi-Aspect Reward Optimization for MLLM Alignment
- Agent-ScanKit: Unraveling Memory and Reasoning of Multimodal Agents via Sensitivity Perturbations
- DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- Latent Visual Reasoning
- Model Correlation Detection via Random Selection Probing
- GeoVLM-R1: Reinforcement Fine-Tuning for Improved Remote Sensing Reasoning
- PIPer: On-Device Environment Setup via Online Reinforcement Learning
- Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling
- Decoupling Reasoning and Perception: An LLM-LMM Framework for Faithful Visual Reasoning
- Unlocking the Essence of Beauty: Advanced Aesthetic Reasoning with Relative-Absolute Policy Optimization
- Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
- Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
- Learning GUI Grounding with Spatial Reasoning from Visual Feedback
- SFT Doesn't Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs
- RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
- DRES: Benchmarking LLMs for Disfluency Removal
- Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions
- OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation
- Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
- Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
- Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding
- Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
- Reinforced Visual Perception with Tools
- Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
- UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning
- UI-Venus Technical Report: Building High-performance UI Agents with RFT
- Diversity First, Quality Later: A Two-Stage Assumption for Language Model Alignment
- We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
- Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
- AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
- Improving Task Diversity in Label Efficient Supervised Finetuning of LLMs
- Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning
- M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
- The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
- CogniSQL-R1-Zero: Lightweight Reinforced Reasoning for Efficient SQL Generation
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
- StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
- WebSailor: Navigating Super-human Reasoning for Web Agent
- Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
- Empowering Small VLMs to Think with Dynamic Memorization and Exploration
- MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
- SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
- Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
- VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?
Related