Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
2025/07/17 by Hao Sun, Mihaela van der Schaar, Sun, Hao +1 · 4 citations
Computer Science · #Speech and dialogue systems
paper · pdf · doi:10.48550/arxiv.2507.13158
Abstract
In the era of Large Language Models (LLMs), alignment has emerged as a fundamental yet challenging problem in the pursuit of more reliable, controllable, and capable machine intelligence. The recent success of reasoning models and conversational AI systems has underscored the critical role of reinforcement learning (RL) in enhancing these systems, driving increased research interest at the intersection of RL and LLM alignment. This paper provides a comprehensive review of recent advances in LLM alignment through the lens of inverse reinforcement learning (IRL), emphasizing the distinctions between RL techniques employed in LLM alignment and those in conventional RL tasks. In particular, we highlight the necessity of constructing neural reward models from human data and discuss the formal and practical implications of this paradigm shift. We begin by introducing fundamental concepts in RL to provide a foundation for readers unfamiliar with the field. We then examine recent advances in this research agenda, discussing key challenges and opportunities in conducting IRL for LLM alignment. Beyond methodological considerations, we explore practical aspects, including datasets, benchmarks, evaluation metrics, infrastructure, and computationally efficient training and inference techniques. Finally, we draw insights from the literature on sparse-reward RL to identify open questions and potential research directions. By synthesizing findings from diverse studies, we aim to provide a structured and critical overview of the field, highlight unresolved challenges, and outline promising future directions for improving LLM alignment through RL and IRL techniques.
Citations
- GRAM: A Generative Foundation Reward Model for Reward Generalization
- Spurious Rewards: Rethinking Training Signals in RLVR
- Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
- OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models
- Reinforcement Learning for Reasoning in Large Language Models with One Training Example
- A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
- Inference-Time Scaling for Generalist Reward Modeling
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Rethinking Diverse Human Preference Learning through Principal Component Analysis
- A Survey of Automatic Prompt Engineering: An Optimization Perspective
- PILAF: Optimal Human Preference Sampling for Reward Modeling
- Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs
- Reviving The Classics: Active Reward Modeling in Large Language Model Alignment
- Reward-Guided Speculative Decoding for Efficient LLM Reasoning
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
- Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes
- Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives
- WorldSimBench: Towards Video Generation Models as World Simulators
- How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective
- PAD: Personalized Alignment of LLMs at Decoding-Time
- Generative Reward Models
- RRM: Robust Reward Model Training Mitigates Reward Hacking
- Generative Verifiers: Reward Modeling as Next-Token Prediction
- Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
- Explaining Length Bias in LLM-Based Preference Evaluations
- Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs
- Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
- A Critical Look At Tokenwise Reward-Guided Text Generation
- ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
- Scalable Ensembling For Mitigating Reward Overoptimisation
- BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling
- Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM Alignment
- Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning
- SimPO: Simple Preference Optimization with a Reference-Free Reward
- DPO Meets PPO: Reinforced Token Optimization for RLHF
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- Overcoming Reward Overoptimization via Adversarial Policy Optimization with Lightweight Uncertainty Estimation
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
- MaxMin-RLHF: Alignment with Diverse Human Preferences
- Active Preference Learning for Large Language Models
- Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts
- A Roadmap to Pluralistic Alignment
- Direct Language Model Alignment from Online AI Feedback
- Personalized Language Modeling from Personalized Human Feedback
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- Dense Reward for Free in Reinforcement Learning from Human Feedback
- ARGS: Alignment as Reward-Guided Search
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
- Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
- Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
- A Survey of Temporal Credit Assignment in Deep Reinforcement Learning
- Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs
- When is Off-Policy Evaluation (Reward Modeling) Useful in Contextual Bandits? A Data-Centric Perspective
- Instruction-Following Evaluation for Large Language Models
- Contrastive Preference Learning: Learning from Human Feedback without RL
- A General Theoretical Paradigm to Understand Learning from Human Preferences
- Privacy in Large Language Models: Attacks, Defenses and Future Directions
- Reward Model Ensembles Help Mitigate Overoptimization
- Large Language Models Cannot Self-Correct Reasoning Yet
- Query-Dependent Prompt Evaluation and Optimization with Offline Inverse RL
- Statistical Rejection Sampling Improves Preference Optimization
- Large Language Models as Optimizers
- Diversifying AI: Towards Creative Chess with AlphaZero
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Secrets of RLHF in Large Language Models Part I: PPO
- Style Over Substance: Evaluation Biases for Large Language Models
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation
- Simple and Controllable Music Generation
- For SALE: State-Action Representation Learning for Deep Reinforcement Learning
- AI Transparency in the Age of LLMs: A Human-Centered Research Roadmap
- Let's Verify Step by Step
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
- Large Language Models are not Fair Evaluators
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Improving Factuality and Reasoning in Language Models through Multiagent Debate
- SLiC-HF: Sequence Likelihood Calibration with Human Feedback
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Automatic Prompt Optimization with "Gradient Descent" and Beam Search
- RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
- RRHF: Rank Responses to Align Language Models with Human Feedback without tears
- GPT-4 Technical Report
- PaLM-E: An Embodied Multimodal Language Model
- Offline Imitation Learning with Suboptimal Demonstrations via Relaxed Distribution Matching
- Adding Conditional Control to Text-to-Image Diffusion Models
- Adding Conditional Control to Text-to-Image Diffusion Models
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
- Affective Coherence Monitoring for Transformer-Based Language Models
- Large Language Models Are Human-Level Prompt Engineers
- When Life Gives You Lemons, Make Cherryade: Converting Feedback from Bad Responses into Good Labels
- Scaling Laws for Reward Model Overoptimization
- Atari-5: Distilling the Arcade Learning Environment down to Five Games
- Decomposed Prompting: A Modular Approach for Solving Complex Tasks
- Prompt-to-Prompt Image Editing with Cross Attention Control
- Emergent Abilities of Large Language Models
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Large Language Models are Zero-Shot Reasoners
- Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
- A Generalist Agent
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Training Compute-Optimal Large Language Models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Training language models to follow instructions with human feedback
- A Survey of Explainable Reinforcement Learning
- Rethinking Goal-conditioned Supervised Learning and Its Connection to Offline RL
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Learning to Repair: Repairing model output errors after deployment using a dynamic memory of feedback
- Learning Long-Term Reward Redistribution via Randomized Return Decomposition
- What Matters for Adversarial Imitation Learning?
- Zero-Shot Text-to-Image Generation
- Offline Learning from Demonstrations and Unlabeled Experience
- f-IRL: Inverse Reinforcement Learning via State Marginal Matching
- DiffWave: A Versatile Diffusion Model for Audio Synthesis
- Learning to summarize from human feedback
- Conservative Q-Learning for Offline Reinforcement Learning
- Strictly Batch Imitation Learning by Energy-based Distribution Matching
- Language Models are Few-Shot Learners
- Scaling Laws for Neural Language Models
- Dota 2 with Large Scale Deep Reinforcement Learning
- Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
- A Divergence Minimization Perspective on Imitation Learning Methods
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Imitation Learning as f-Divergence Minimization
- Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations
- Soft Actor-Critic Algorithms and Applications
- Off-Policy Deep Reinforcement Learning without Exploration
- Sparse Attentive Backtracking: Temporal CreditAssignment Through Reminding
- RUDDER: Return Decomposition for Delayed Rewards
- Addressing Function Approximation Error in Actor-Critic Methods
- DeepMind Control Suite
- Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
- Learning Robust Rewards with Adversarial Inverse Reinforcement Learning
- Overcoming Exploration in Reinforcement Learning with Demonstrations
- Proximal Policy Optimization Algorithms
- Hindsight Experience Replay
- Deep reinforcement learning from human preferences
- Attention Is All You Need
- Curiosity-driven Exploration by Self-supervised Prediction
- Generative Adversarial Imitation Learning
- f-GAN: Training Generative Neural Samplers using Variational Divergence Minimization
- High-Dimensional Continuous Control Using Generalized Advantage\n Estimation
- Playing Atari with Deep Reinforcement Learning
- A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
- Bayesian Experimental Design: A Review
- RANK ANALYSIS OF INCOMPLETE BLOCK DESIGNS
- Retrieval Augmented Thought Process for Private Data Handling in Healthcare
- OpenAI o1 System Card
- EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
- A Survey on Explainable Reinforcement Learning: Concepts, Algorithms, Challenges
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Cited by
Related