On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
2025/08/07 by Wu, Yongliang, Zhou, Yizhou, Ziheng, Zhou +7 · 35 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
paper · doi:10.48550/arxiv.2508.05629
Abstract
We present a simple yet theoretically motivated improvement to Supervised Fine-Tuning (SFT) for the Large Language Model (LLM), addressing its limited generalization compared to reinforcement learning (RL). Through mathematical analysis, we reveal that standard SFT gradients implicitly encode a problematic reward structure that may severely restrict the generalization capabilities of model. To rectify this, we propose Dynamic Fine-Tuning (DFT), stabilizing gradient updates for each token by dynamically rescaling the objective function with the probability of this token. Remarkably, this single-line code change significantly outperforms standard SFT across multiple challenging benchmarks and base models, demonstrating greatly improved generalization. Additionally, our approach shows competitive results in offline RL settings, offering an effective yet simpler alternative. This work bridges theoretical insight and practical solutions, substantially advancing SFT performance. The code will be available at https://github.com/yongliang-wu/DFT.
Citations
- Hierarchical Semantic Alignment for Image Clustering
- Enhancing CLIP Robustness via Cross-Modality Alignment
- MM-HELIX: Boosting Multimodal Long-Chain Reflective Reasoning with Holistic Platform and Adaptive Hybrid Policy Optimization
- Real-Time Motion-Controllable Autoregressive Video Diffusion
- VideoNSA: Native Sparse Attention Scales Video Understanding
- ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
- HiFi-Mamba: Dual-Stream W-Laplacian Enhanced Mamba for High-Fidelity MRI Reconstruction
- TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
- Group Sequence Policy Optimization
- Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
- The Hidden Costs of AI: A Review of Energy, E-Waste, and Inequality in Model Development
- Dynamic Multimodal Prototype Learning in Vision-Language Models
- AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
- Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
- Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
- Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
- VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism
- WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
- Bridging Supervised Learning and Reinforcement Learning in Math Reasoning
- UFT: Unifying Supervised and Reinforcement Fine-Tuning
- Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization
- Learning to Reason under Off-Policy Guidance
- Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark
- Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
- Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
- SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion Models
- An Empirical Study on Eliciting and Improving R1-like Reasoning Models
- All Roads Lead to Likelihood: The Value of Reinforcement Learning in Fine-Tuning
- RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- Qwen2.5 Technical Report
- Enhancing Zero-Shot Vision Models by Label-Free Prompt Distribution Learning and Bias Correcting
- Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
- Learning from negative feedback, or positive feedback or both
- AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
- SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning
- The Llama 3 Herd of Models
- Selective Vision-Language Subspace Projection for Few-shot CLIP
- VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
- We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
- SEED: Accelerating Reasoning Tree Construction via Scheduled Speculative Decoding
- MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
- Green AI: Exploring Carbon Footprints, Mitigation Strategies, and Trade Offs in Large Language Model Training
- Boosting Few-Shot Learning via Attentive Feature Regularization
- MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
- LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
- Causal Prompting: Debiasing Large Language Model Prompting based on Front-Door Adjustment
- Causal Walk: Debiasing Multi-Hop Fact Verification with Front-Door Adjustment
- Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- Large Language Models for Mathematical Reasoning: Progresses and Challenges
- UltraFeedback: Boosting Language Models with Scaled AI Feedback
- Instruction Tuning for Large Language Models: A Survey
- MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- MixCE: Training Autoregressive Language Models by Mixing Forward and Reverse Cross-Entropies
- Hi-ResNet: Edge Detail Enhancement for High-Resolution Remote Sensing Segmentation
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
- RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
- Scaling Instruction-Finetuned Language Models
- Solving Quantitative Reasoning Problems with Language Models
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Training language models to follow instructions with human feedback
- Multitask Prompted Training Enables Zero-Shot Task Generalization
- Finetuned Language Models Are Zero-Shot Learners
- What Matters in Learning from Offline Human Demonstrations for Robot Manipulation
- Evaluating Large Language Models Trained on Code
- Measuring Mathematical Problem Solving With the MATH Dataset
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Energy and Policy Considerations for Deep Learning in NLP
- A Mean Field Theory of Batch Normalization
- Proximal Policy Optimization Algorithms
- Deep reinforcement learning from human preferences
- On the difficulty of training Recurrent Neural Networks
- A more robust boosting algorithm
- Text Generation by Learning from Demonstrations
Cited by
Related