Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
2025/09/20 by Keliang Liu, Dingkang Yang, Liu, Keliang +15 · 5 citations
Mathematics · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Modeling, Simulation, and Optimization
paper · pdf · doi:10.48550/arxiv.2509.16679
openalex publication_date 2025/09/20 · openalex created_date 2025/10/16 · openalex updated_date 2026/07/28
Abstract
In recent years, training methods centered on Reinforcement Learning (RL) have markedly enhanced the reasoning and alignment performance of Large Language Models (LLMs), particularly in understanding human intents, following user instructions, and bolstering inferential strength. Although existing surveys offer overviews of RL augmented LLMs, their scope is often limited, failing to provide a comprehensive summary of how RL operates across the full lifecycle of LLMs. We systematically review the theoretical and practical advancements whereby RL empowers LLMs, especially Reinforcement Learning with Verifiable Rewards (RLVR). First, we briefly introduce the basic theory of RL. Second, we thoroughly detail application strategies for RL across various phases of the LLM lifecycle, including pre-training, alignment fine-tuning, and reinforced reasoning. In particular, we emphasize that RL methods in the reinforced reasoning phase serve as a pivotal driving force for advancing model reasoning to its limits. Next, we collate existing datasets and evaluation benchmarks currently used for RL fine-tuning, spanning human-annotated datasets, AI-assisted preference data, and program-verification-style corpora. Subsequently, we review the mainstream open-source tools and training frameworks available, providing clear practical references for subsequent research. Finally, we analyse the future challenges and trends in the field of RL-enhanced LLMs. This survey aims to present researchers and practitioners with the latest developments and frontier trends at the intersection of RL and LLMs, with the goal of fostering the evolution of LLMs that are more intelligent, generalizable, and secure.
Citations
- A Survey of Reinforcement Learning for Large Reasoning Models
- Adaptive Reinforcement for Open-ended Medical Reasoning via Semantic-Guided Reward Collapse Mitigation
- Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
- 3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
- Graph-R1: Towards Agentic GraphRAG Framework via End-to-end Reinforcement Learning
- Post-Training Large Language Models via Reinforcement Learning from Self-Feedback
- Group Sequence Policy Optimization
- DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training
- Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
- Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
- MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
- OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
- Thought Anchors: Which LLM Reasoning Steps Matter?
- No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
- VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
- AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
- Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
- TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
- ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
- SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
- τ2-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
- Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
- RewardAnything: Generalizable Principle-Following Reward Models
- Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
- The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
- SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
- SuperRL: Reinforcement Learning with Supervision to Boost Language Model Reasoning
- GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
- AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time
- AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- Reinforcing General Reasoning without Verifiers
- SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
- CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric Reward
- Learning to Reason without External Rewards
- Adaptive Deep Reasoning: Triggering Deep Thinking When Needed
- More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
- SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward
- LLM-Powered AI Agent Systems and Their Applications in Industry
- RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
- When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning
- RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning
- Reward Reasoning Model
- Think Only When You Need with Large Hybrid-Reasoning Models
- Thinkless: LLM Learns When to Think
- RL in Name Only? Analyzing the Structural Assumptions in RL post-training for LLMs
- ToTRL: Unlock LLM Tree-of-Thoughts Reasoning Potential through Puzzles Solving
- AdaptThink: Reasoning Models Can Learn When to Think
- SLOT: Sample-specific Language Model Optimization at Test-time
- Visual Planning: Let's Think Only with Images
- Qwen3 Technical Report
- OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
- DanceGRPO: Unleashing GRPO on Visual Generation
- Absolute Zero: Reinforced Self-play Reasoning with Zero Data
- Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
- X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
- RM-R1: Reward Modeling as Reasoning
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
- ShorterBetter: Guiding Reasoning Models to Find Optimal Inference Length for Efficient Reasoning
- SWE-smith: Scaling Data for Software Engineering Agents
- PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
- Learning to Reason under Off-Policy Guidance
- SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Compile Scene Graphs with Reinforcement Learning
- Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
- Reasoning Models Can Be Effective Without Thinking
- TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning
- A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems
- SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
- Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use
- Rethinking Reflection in Pre-Training
- Inference-Time Scaling for Generalist Reward Modeling
- SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
- On Data Synthesis and Post-training for Visual Abstract Reasoning
- GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
- Improved Visual-Spatial Reasoning via R1-Zero-Like Training
- Boosting MLLM Reasoning with Text-Debiased Hint-GRPO
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
- Video-R1: Reinforcing Video Reasoning in MLLMs
- Think or Not Think: A Study of Explicit Thinking in Rule-Based Visual Reinforcement Fine-Tuning
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- DeepPerception: Advancing R1-like Cognitive Visual Perception in MLLMs for Knowledge-Intensive Visual Grounding
- R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
- L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
- Visual-RFT: Visual Reinforcement Fine-Tuning
- Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?
- BIG-Bench Extra Hard
- Reward Shaping to Mitigate Reward Hacking in RLHF
- Scalable Best-of-N Selection for Large Language Models via Self-Certainty
- ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
- The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking
- MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- Reasoning Language Models: A Blueprint
- Reinforcement Learning Enhanced LLMs: A Survey
- Improving Factuality in Large Language Models via Decoding-Time Hallucinatory and Truthful Comparators
- A Comprehensive Survey of LLM Alignment Techniques: RLHF, RLAIF, PPO, DPO and More
- Large Vision-Language Models as Emotion Recognizers in Context Awareness
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric Applications
- LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence
- AGILE: A Novel Reinforcement Learning Framework of LLM Agents
- Efficiency in Focus: LayerNorm as a Catalyst for Fine-tuning Medical Visual Language Pre-trained Models
- RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- ORPO: Monolithic Preference Optimization without Reference Model
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
- InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- A Minimaximalist Approach to Reinforcement Learning from Human Feedback
- A Survey of Reinforcement Learning from Human Feedback
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Instruction-Following Evaluation for Large Language Models
- A General Theoretical Paradigm to Understand Learning from Human Preferences
- Large Language Models Cannot Self-Correct Reasoning Yet
- Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
- Understanding Catastrophic Forgetting in Language Models via Implicit Inference
- DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- TheoremQA: A Theorem-driven Question Answering dataset
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
- A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
- Affective Coherence Monitoring for Transformer-Based Language Models
- Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
- Solving Quantitative Reasoning Problems with Language Models
- On Reinforcement Learning and Distribution Matching for Fine-Tuning Language Models with no Catastrophic Forgetting
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Training language models to follow instructions with human feedback
- Ethical and social risks of harm from Language Models
- A General Language Assistant as a Laboratory for Alignment
- Training Verifiers to Solve Math Word Problems
- On the Opportunities and Risks of Foundation Models
- GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow
- Alignment of Language Agents
- On the Dangers of Stochastic Parrots
- Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models
- RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language\n Models
- Measuring Massive Multitask Language Understanding
- Proximal Policy Optimization Algorithms
- Trust Region Policy Optimization
- The Invisible Leash: Why RLVR May or May Not Escape Its Origin
- ACEBench: Who Wins the Match Point in Tool Usage?
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
- AI Benchmark Half-Life in Recursive Corpora: A Theory of Validity Decay under Semantic Leakage and Regeneration
- OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Cited by
Related