A Survey on Progress in LLM Alignment from the Perspective of Reward Design
2025/05/05 by Min Ji, Yonghui Wu, Ji, Miaomiao +11 · 10 citations
Environmental Science · #Educational Reforms and Innovations
paper · pdf · doi:10.48550/arxiv.2505.02666
Abstract
Reward design plays a pivotal role in aligning large language models (LLMs) with human values, serving as the bridge between feedback signals and model optimization. This survey provides a structured organization of reward modeling and addresses three key aspects: mathematical formulation, construction practices, and interaction with optimization paradigms. Building on this, it develops a macro-level taxonomy that characterizes reward mechanisms along complementary dimensions, thereby offering both conceptual clarity and practical guidance for alignment research. The progression of LLM alignment can be understood as a continuous refinement of reward design strategies, with recent developments highlighting paradigm shifts from reinforcement learning (RL)-based to RL-free optimization and from single-task to multi-objective and complex settings.
Citations
- AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
- AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models
- Discriminative Policy Optimization for Token-Level Reward Models
- MOSLIM:Align with diverse preferences in prompts through reward classification
- BLEUBERI: BLEU is a surprisingly effective reward for instruction following
- A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future
- Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
- Sentence-level Reward Model can Generalize Better for Aligning LLM from Human Preference
- Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems
- VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare
- Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
- Approximated Variational Bayesian Inverse Reinforcement Learning for Large Language Model Alignment
- Constrain Alignment with Sparse Autoencoders
- Rule Based Rewards for Language Model Safety
- Optimal Design for Reward Modeling in RLHF
- A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications
- TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights
- Generative Reward Models
- Generative Verifiers: Reward Modeling as Next-Token Prediction
- Selective Preference Optimization via Token-Level Reward Function Estimation
- TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
- A Comprehensive Survey of LLM Alignment Techniques: RLHF, RLAIF, PPO, DPO and More
- Visual Prompt Selection for In-Context Learning Segmentation
- β-DPO: Direct Preference Optimization with Dynamic β
- Sub-SA: Strengthen In-context Learning via Submodular Selective Annotation
- HAF-RM: A Hybrid Alignment Framework for Reward Model Training
- Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
- ICLEval: Evaluating In-Context Learning Ability of Large Language Models
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
- A Survey on Human Preference Learning for Large Language Models
- OPTune: Efficient Online Preference Tuning
- Aligning Large Language Models via Fine-grained Supervision
- SimPO: Simple Preference Optimization with a Reference-Free Reward
- SPO: Multi-Dimensional Preference Sequential Alignment With Implicit Reward Modeling
- RLHF Workflow: From Reward Modeling to Online RLHF
- Iterative Reasoning Preference Optimization
- DPO Meets PPO: Reinforced Token Optimization for RLHF
- Token-level Direct Preference Optimization
- Disentangling Length from Quality in Direct Preference Optimization
- Large Language Models for Education: A Survey and Outlook
- ORPO: Monolithic Preference Optimization without Reference Model
- Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards
- ALaRM: Align Language Models via Hierarchical Rewards Modeling
- On the Essence and Prospect: An Investigation of Alignment Approaches for Big Models
- Unsupervised Zero-Shot Reinforcement Learning via Functional Reward Encodings
- Generalizing Reward Modeling for Out-of-Distribution Preference Learning
- Reward Generalization in RLHF: A Topological Perspective
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- Dense Reward for Free in Reinforcement Learning from Human Feedback
- Self-Rewarding Language Models
- Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation
- Improving Large Language Models via Fine-grained Reinforcement Learning with Minimum Editing Constraint
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
- Aligning Large Language Models with Human Preferences through Representation Engineering
- Reasons to Reject? Aligning Language Models with Judgments
- Prompt Optimization via Adversarial In-Context Learning
- Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game
- Beyond Imitation: Leveraging Fine-grained Quality Signals for Alignment
- Instruct Me More! Random Prompting for Visual In-Context Learning
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- Improving Generalization of Alignment with Human Preferences through Group Invariant Learning
- A General Theoretical Paradigm to Understand Learning from Human Preferences
- ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
- Improving Factual Consistency for Knowledge-Grounded Dialogue Systems via Knowledge Enhancement and Alignment
- SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF
- Confronting Reward Model Overoptimization with Constrained RLHF
- Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
- Reward Model Ensembles Help Mitigate Overoptimization
- Tool-Augmented Reward Modeling
- Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment
- Large Language Model Alignment: A Survey
- Aligning Large Multimodal Models with Factually Augmented RLHF
- SYNDICOM: Improving Conversational Commonsense with Error-Injection and Natural Language Feedback
- Efficient RLHF: Reducing the Memory Usage of PPO
- Shepherd: A Critic for Language Model Generation
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Secrets of RLHF in Large Language Models Part I: PPO
- Lost in the Middle: How Language Models Use Long Contexts
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
- Fine-Grained Human Feedback Gives Better Rewards for Language Model Training
- Preference-grounded Token-level Guidance for Language Model Fine-tuning
- Let's Verify Step by Step
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Aligning Large Language Models through Synthetic Feedback
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
- LIMA: Less Is More for Alignment
- RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs
- RRHF: Rank Responses to Align Language Models with Human Feedback without tears
- Self-Refine: Iterative Refinement with Self-Feedback
- GPT-4 Technical Report
- Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Compositional Exemplars for In-context Learning
- Extracting Training Data from Diffusion Models
- MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning
- Self-Instruct: Aligning Language Models with Self-Generated Instructions
- Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor
- Affective Coherence Monitoring for Transformer-Based Language Models
- Large Language Models Are Human-Level Prompt Engineers
- Scaling Laws for Reward Model Overoptimization
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
- Defining and Characterizing Reward Hacking
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Training language models to follow instructions with human feedback
- Survey of Hallucination in Natural Language Generation
- WebGPT: Browser-assisted question-answering with human feedback
- A General Language Assistant as a Laboratory for Alignment
- On the Opportunities and Risks of Foundation Models
- Multimodal Few-Shot Learning with Frozen Language Models
- Societal Biases in Language Generation: Progress and Challenges
- On the Dangers of Stochastic Parrots
- RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
- Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO
- Artificial Intelligence, Values and Alignment
- Fine-Tuning Language Models from Human Preferences
- On the Weaknesses of Reinforcement Learning for Neural Machine Translation
- Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
- Proximal Policy Optimization Algorithms
- Deep reinforcement learning from human preferences
- Attention Is All You Need
- Concrete Problems in AI Safety
- The Analysis of Permutations
- RANK ANALYSIS OF INCOMPLETE BLOCK DESIGNS
- UltraFeedback: Boosting Language Models with Scaled AI Feedback
- SALMON: Self-Alignment with Instructable Reward Models
- Aligning Large Language Models with Human: A Survey
Cited by
Related