SimPO: Simple Preference Optimization with a Reference-Free Reward
2024/05/23 by Yu Meng, Meng, Yu, Mengzhou Xia +3 · 174 citations
Computer Science · Decision Sciences · Engineering · #Computation and Language (cs.CL) #Constraint Satisfaction and Optimization #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multi-Criteria Decision Making #Transportation and Mobility Innovations
paper · pdf · doi:10.48550/arxiv.2405.14734
openalex publication_date 2024/05/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Direct Preference Optimization (DPO) is a widely used offline preference optimization algorithm that reparameterizes reward functions in reinforcement learning from human feedback (RLHF) to enhance simplicity and training stability. In this work, we propose SimPO, a simpler yet more effective approach. The effectiveness of SimPO is attributed to a key design: using the average log probability of a sequence as the implicit reward. This reward formulation better aligns with model generation and eliminates the need for a reference model, making it more compute and memory efficient. Additionally, we introduce a target reward margin to the Bradley-Terry objective to encourage a larger margin between the winning and losing responses, further improving the algorithm's performance. We compare SimPO to DPO and its latest variants across various state-of-the-art training setups, including both base and instruction-tuned models such as Mistral, Llama 3, and Gemma 2. We evaluate on extensive chat-based evaluation benchmarks, including AlpacaEval 2, MT-Bench, and Arena-Hard. Our results demonstrate that SimPO consistently and significantly outperforms existing approaches without substantially increasing response length. Specifically, SimPO outperforms DPO by up to 6.4 points on AlpacaEval 2 and by up to 7.5 points on Arena-Hard. Our top-performing model, built on Gemma-2-9B-it, achieves a 72.4% length-controlled win rate on AlpacaEval 2, a 59.1% win rate on Arena-Hard, and ranks 1st on Chatbot Arena among <10B models with real user votes.
Cited by
- Alpha-R1: Alpha Screening with LLM Reasoning via Reinforcement Learning
- APO: Alpha-Divergence Preference Optimization
- Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization
- FlashEvaluator: Expanding Search Space with Parallel Sequence-Level Evaluation
- Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
- GoldenFuzz: Generative Golden Reference Hardware Fuzzing
- ORPR: An OR-Guided Pretrain-then-Reinforce Learning Model for Inventory Management
- Training LLMs with LogicReward for Faithful and Rigorous Reasoning
- Structure-Aware Antibody Design with Affinity-Optimized Inverse Folding
- TakeAD: Preference-based Post-optimization for End-to-end Autonomous Driving with Expert Takeover Data
- Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
- Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
- AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
- Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
- SpeakRL: Synergizing Reasoning, Speaking, and Acting in Language Models with Reinforcement Learning
- LLM-Auction: Generative Auction towards LLM-Native Advertising
- Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
- SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language Models
- When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
- Policy-based Sentence Simplification: Replacing Parallel Corpora with LLM-as-a-Judge
- Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts
- What Is Preference Optimization Doing, How and Why?
- When Human Preferences Flip: An Instance-Dependent Robust Loss for RLHF
- Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization
- Real-Time Long Horizon Air Quality Forecasting via Group-Relative Policy Optimization
- Lightweight Model Editing for LLMs to Correct Deprecated API Recommendations
- Beyond Reward Margin: Rethinking and Resolving Likelihood Displacement in Diffusion Models via Video Generation
- Multimodal Large Language Models with Adaptive Preference Optimization for Sequential Recommendation
- UFO: Unfair-to-Fair Evolving Mitigates Unfairness in LLM-based Recommender Systems via Self-Play Fine-tuning
- Incorporating Self-Rewriting into Large Language Model Reasoning Reinforcement
- TB or Not TB: Coverage-Driven Direct Preference Optimization for Verilog Stimulus Generation
- Bootstrapping LLMs via Preference-Based Policy Optimization
- AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment
- LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
- Convergence and Stability Analysis of Self-Consuming Generative Models with Heterogeneous Human Curation
- The Path Not Taken: RLVR Provably Learns Off the Principals
- PC-Diffusion: Aligning Diffusion Models with Human Preferences via Preference Classifier
- CAPO: Confidence Aware Preference Optimization Learning for Multilingual Preferences
- SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization
- Directional-Clamp PPO
- Inference-Time Personalized Alignment with a Few User Preference Queries
- Value Drifts: Tracing Value Alignment During LLM Post-Training
- BOTS: A Unified Framework for Bayesian Online Task Selection in LLM Reinforcement Finetuning
- Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback
- Mediocrity is the key for LLM as a Judge Anchor Selection
- Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?
- Sharpness-Guided Group Relative Policy Optimization via Probability Shaping
- Semi-Supervised Preference Optimization with Limited Feedback
- GIFT: Group-relative Implicit Fine Tuning Integrates GRPO with DPO and UNA
- POPI: Personalizing LLMs via Optimized Natural Language Preference Inference
- MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding
- Offline Preference Optimization via Maximum Marginal Likelihood Estimation
- Aligning Diffusion Language Models via Unpaired Preference Optimization
- Learning "Partner-Aware" Collaborators in Multi-Party Collaboration
- Beyond Pairwise: Empowering LLM Alignment With Ranked Choice Modeling
- Why DPO is a Misspecified Estimator and How to Fix It
- SynCast: Synergizing Contradictions in Precipitation Nowcasting via Diffusion Sequential Preference Optimization
- Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
- Rectifying Shortcut Behaviors in Preference-based Reward Learning
- Towards Flash Thinking via Decoupled Advantage Policy Optimization
- Rethinking On-policy Optimization for Query Augmentation
- On Efficiency-Effectiveness Trade-off of Diffusion-based Recommenders
- How Well Can Preference Optimization Generalize Under Noisy Feedback?
- Holdout-Loss-Based Data Selection for LLM Finetuning via In-Context Learning
- M2PO: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- Towards Understanding Valuable Preference Data for Large Language Model Alignment
- Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
- On the Role of Preference Variance in Preference Optimization
- Reinforced Preference Optimization for Recommendation
- Enhancing LLM Reasoning via Non-Human-Like Reasoning Path Preference Optimization
- Beyond the limitation of a single query: Train your LLM for query expansion with Reinforcement Learning
- RIPRAG: Hack a Black-box Retrieval-Augmented Generation Question-Answering System with Reinforcement Learning
- CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
- Contrastive Weak-to-strong Generalization
- Can Speech LLMs Think while Listening?
- Tool-Augmented Policy Optimization: Synergizing Reasoning and Adaptive Tool Use with Reinforcement Learning
- Predictive Preference Learning from Human Interventions
- Aligning Large Language Models via Fully Self-Synthetic Data
- Towards Better Optimization For Listwise Preference in Diffusion Models
- On the Role of Difficult Prompts in Self-Play Preference Optimization
- Provably Mitigating Corruption, Overoptimization, and Verbosity Simultaneously in Offline and Online RLHF/DPO Alignment
- Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization
- From Noisy Traces to Stable Gradients: Bias-Variance Optimized Preference Optimization for Aligning Large Reasoning Models
- Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning
- Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents
- Fine-Tuning Diffusion Models via Intermediate Distribution Shaping
- It Takes Two: Your GRPO Is Secretly DPO
- Rethinking Reward Models for Multi-Domain Test-Time Scaling
- AP2O-Coder: Adaptively Progressive Preference Optimization for Reducing Compilation and Runtime Errors in LLM-Generated Code
- MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation
- Scalable and Robust LLM Unlearning by Correcting Responses with Retrieved Exclusions
- Distillation of Large Language Models via Concrete Score Matching
- Improving Sampling Efficiency in RLVR through Adaptive Rollout and Response Reuse
- Probing the Limits of Stylistic Alignment in Vision-Language Models
- Can Molecular Foundation Models Know What They Don't Know? A Simple Remedy with Preference Optimization
- Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling
- SeaPO: Strategic Error Amplification for Robust Preference Optimization of Large Language Models
- Humanline: Online Alignment as Perceptual Loss
- Model Correlation Detection via Random Selection Probing
- RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM Alignment
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
- Toward Preference-aligned Large Language Models via Residual-based Model Steering
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
- Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization
- Test-Time Policy Adaptation for Enhanced Multi-Turn Interactions with LLMs
- Multiplayer Nash Preference Optimization
- OFMU: Optimization-Driven Framework for Machine Unlearning
- ResT: Reshaping Token-Level Policy Gradients for Tool-Use Large Language Models
- RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards
- Failure Modes of Maximum Entropy RLHF
- Future Policy Aware Preference Learning for Mathematical Reasoning
- Embedding Domain Knowledge for Large Language Models via Reinforcement Learning from Augmented Generation
- UserRL: Training Interactive User-Centric Agent via Reinforcement Learning
- MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
- ToolRec: Calibrated Preference Alignment for Query Recommendation in On-Device Assistants
- References Improve LLM Alignment in Non-Verifiable Domains
- Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model Diffing
- OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation
- DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving
- Preference Distillation via Value based Reinforcement Learning
- From Uniform to Heterogeneous: Tailoring Policy Optimization to Every Token's Nature
- Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
- Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment
- Relevance to Utility: Process-Supervised Rewrite for RAG
- GPO: Learning from Critical Steps to Improve LLM Reasoning
- Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
- The Anatomy of Alignment: Decomposing Preference Optimization by Steering Sparse Features
- Tool-R1: Sample-Efficient Reinforcement Learning for Agentic Tool Use
- Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization
- When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models
- What Matters in Data for DPO?
- When Are Two RLHF Objectives the Same?
- CM-Align: Consistency-based Multilingual Alignment for Large Language Models
- Generative quantum eigensolver with constrained circuit-cutting overhead
- EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention
- Let's Roleplay: Examining LLM Alignment in Collaborative Dialogues
- Finetuning LLMs for Human Behavior Prediction in Social Science Experiments
- Icon2: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation
- What-If Analysis of Large Language Models: Explore the Game World Using Proactive Thinking
- Delta Activations: A Representation for Finetuned Large Language Models
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
- Improving Large Vision and Language Models by Learning from a Panel of Peers
- MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
- Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models
- HEAL: A Hypothesis-Based Preference-Aware Analysis Framework
- FormaRL: Enhancing Autoformalization with no Labeled Data
- Harnessing Rule-Based Reinforcement Learning for Enhanced Grammatical Error Correction
- Better Language Model-Based Judging Reward Modeling through Scaling Comprehension Boundaries
- Linear Preference Optimization: Decoupled Gradient Control via Absolute Regularization
- ChronoLLM: Customizing Language Models for Physics-Based Simulation Code Generation
- Toward Better EHR Reasoning in LLMs: Reinforcement Learning with Expert Attention Guidance
- Fusing Rewards and Preferences in Reinforcement Learning
- FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation
- Diversity First, Quality Later: A Two-Stage Assumption for Language Model Alignment
- Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization
- On Negative-aware Preference Optimization for Recommendation
- ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
- STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision
- Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization
- Data Selection for LLM Alignment Using Fine-Grained Preferences
- Sample-efficient LLM Optimization with Reset Replay
- P-Aligner: Enabling Pre-Alignment of Language Models via Principled Instruction Synthesis
- Forgetting: A New Mechanism Towards Better Large Language Model Fine-tuning
- Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
- V.I.P. : Iterative Online Preference Distillation for Efficient Video Diffusion Models
- Uni-Layout: Integrating Human Feedback in Unified Layout Generation and Evaluation
- Don't Overthink It: A Survey of Efficient R1-style Large Reasoning Models
- IMU: Influence-guided Machine Unlearning
- Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models
- Repair-R1: Better Test Before Repair
Related