RL makes MLLMs see better than SFT
2025/10/18 by Junha Song, Song, Junha, Sangdoo Yun +7
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2510.16333
openalex publication_date 2025/10/18 · openalex created_date 2025/10/22 · openalex updated_date 2026/07/28
Abstract
A dominant assumption in Multimodal Language Model (MLLM) research is that its performance is largely inherited from the LLM backbone, given its immense parameter scale and remarkable capabilities. This has created a void in the understanding of the vision encoder, which determines how MLLMs perceive images. The recent shift in MLLM training paradigms, from Supervised Finetuning (SFT) to Reinforcement Learning (RL), magnifies this oversight-namely, the significant lack of analysis on how such training reshapes the vision encoder as well as the MLLM. To address this, we first investigate the impact of training strategies on MLLMs, where RL shows a clear advantage over SFT in strongly vision-related VQA benchmarks. Motivated by this, we conduct a critical yet under-explored analysis of the vision encoder of MLLMs through diverse and in-depth experiments, ranging from ImageNet classification and segmentation to gradient visualization. Our results demonstrate that MLLM's post-training strategy (i.e., SFT or RL) not only leads to distinct outcomes on MLLM downstream tasks, but also fundamentally reshapes MLLM's underlying visual representations. Specifically, the key finding of our study is that RL produces stronger and precisely localized visual representations compared to SFT, boosting the ability of the vision encoder for MLLM. We then reframe our findings into a simple recipe for building strong vision encoders for MLLMs, Preference-Instructed Vision OpTimization (PIVOT). When integrated into MLLMs, a PIVOT-trained vision encoder outperforms even larger and more heavily-trained counterparts, despite requiring less than 1% of the computational cost of standard vision pretraining. This result opens an effective and efficient path for advancing the vision backbones of MLLMs. Project page available at https://june-page.github.io/pivot/
Citations
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- LPOI: Listwise Preference Optimization for Vision Language Models
- Qwen3 Technical Report
- LongPerceptualThoughts: Distilling System-2 Reasoning for System-1 Perception
- Perception Encoder: The best visual embeddings are not at the output of the network
- The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
- Scaling Laws for Native Multimodal Models
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- SmolVLM: Redefining small and efficient multimodal models
- Scaling Language-Free Visual Representation Learning
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2.5-VL Technical Report
- CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key
- Qwen2.5 Technical Report
- VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
- Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
- Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
- V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization
- NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
- Locality Alignment Improves Vision-Language Models
- LLaVA-Critic: Learning to Evaluate Multimodal Models
- MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
- LLaVA-OneVision: Easy Visual Task Transfer
- The Llama 3 Herd of Models
- LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- mDPO: Conditional Preference Optimization for Multimodal Large Language Models
- Benchmarking and Improving Detail Image Caption
- RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness
- Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
- The Platonic Representation Hypothesis
- BRAVE: Broadening the visual encoding of vision-language models
- Gemma: Open Models Based on Gemini Research and Technology
- ORPO: Monolithic Preference Optimization without Reference Model
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
- SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
- HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
- A General Theoretical Paradigm to Understand Learning from Human Preferences
- Improved Baselines with Visual Instruction Tuning
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Qwen Technical Report
- Aligning Large Multimodal Models with Factually Augmented RLHF
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Planting a SEED of Vision in Large Language Model
- MMBench: Is Your Multi-modal Model an All-around Player?
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Evaluating Object Hallucination in Large Vision-Language Models
- Visual Instruction Tuning
- DINOv2: Learning Robust Visual Features without Supervision
- Sigmoid Loss for Language Image Pre-Training
- EVA-CLIP: Improved Training Techniques for CLIP at Scale
- GPT-4 Technical Report
- PaLM-E: An Embodied Multimodal Language Model
- LLaMA: Open and Efficient Foundation Language Models
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- Training language models to follow instructions with human feedback
- ClipCap: CLIP Prefix for Image Captioning
- Masked Autoencoders Are Scalable Vision Learners
- Training Verifiers to Solve Math Word Problems
- Perceiver IO: A General Architecture for Structured Inputs & Outputs
- BEiT: BERT Pre-Training of Image Transformers
- Emerging Properties in Self-Supervised Vision Transformers
- Emerging Properties in Self-Supervised Vision Transformers
- Learning Transferable Visual Models From Natural Language Supervision
- Exploring Simple Siamese Representation Learning
- Exploring Simple Siamese Representation Learning
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Measuring Massive Multitask Language Understanding
- DocVQA: A Dataset for VQA on Document Images
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- Language Models are Few-Shot Learners
- A Simple Framework for Contrastive Learning of Visual Representations
- Momentum Contrast for Unsupervised Visual Representation Learning
- Momentum Contrast for Unsupervised Visual Representation Learning
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language\n Generation, Translation, and Comprehension
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- XLNet: Generalized Autoregressive Pretraining for Language Understanding
- HellaSwag: Can a Machine Really Finish Your Sentence?
- Towards VQA Models That Can Read
- Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
- Proximal Policy Optimization Algorithms
- Deep reinforcement learning from human preferences
- Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization
- Microsoft COCO Captions: Data Collection and Evaluation Server
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Related