Process Reinforcement through Implicit Rewards
2025/02/03 by Ganqu Cui, Cui, Ganqu, Lifan Yuan +45 · 98 citations
Computer Science · #Reinforcement Learning in Robotics
paper · pdf · doi:10.48550/arxiv.2502.01456
Abstract
Dense process rewards have proven a more effective alternative to the sparse outcome-level rewards in the inference-time scaling of large language models (LLMs), particularly in tasks requiring complex multi-step reasoning. While dense rewards also offer an appealing choice for the reinforcement learning (RL) of LLMs since their fine-grained rewards have the potential to address some inherent issues of outcome rewards, such as training efficiency and credit assignment, this potential remains largely unrealized. This can be primarily attributed to the challenges of training process reward models (PRMs) online, where collecting high-quality process labels is prohibitively expensive, making them particularly vulnerable to reward hacking. To address these challenges, we propose PRIME (Process Reinforcement through IMplicit rEwards), which enables online PRM updates using only policy rollouts and outcome labels through implict process rewards. PRIME combines well with various advantage functions and forgoes the dedicated reward model training phrase that existing approaches require, substantially reducing the development overhead. We demonstrate PRIME's effectiveness on competitional math and coding. Starting from Qwen2.5-Math-7B-Base, PRIME achieves a 15.1% average improvement across several key reasoning benchmarks over the SFT model. Notably, our resulting model, Eurus-2-7B-PRIME, surpasses Qwen2.5-Math-7B-Instruct on seven reasoning benchmarks with 10% of its training data.
Cited by
- Reinforcement Learning via Self-Distillation
- OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
- LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
- A Comedy of Estimators: On KL Regularization in RL Training of LLMs
- AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
- Trust-Region Adaptive Policy Optimization
- Stepwise Think-Critique: A Unified Framework for Robust and Interpretable LLM Reasoning
- TraPO: A Semi-Supervised Reinforcement Learning Framework for Boosting LLM Reasoning
- GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- Coupled Variational Reinforcement Learning for Language Model General Reasoning
- No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
- SPARK: Stepwise Process-Aware Rewards for Reference-Free Reinforcement Learning
- Adversarial Training for Process Reward Models
- Reinforcing Action Policies by Prophesying
- Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
- P1: Mastering Physics Olympiads with Reinforcement Learning
- Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning
- Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling
- RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
- Extending RLVR to Open-Ended Tasks via Verifiable Multiple-Choice Reformulation
- RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
- Weak-to-Strong On-Policy Distillation
- Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?
- Reasoning-Aware GRPO using Process Mining
- Scheduling Your LLM Reinforcement Learning with Reasoning Trees
- Smaller Models, Smarter Rewards: A Two-Sided Approach to Process and Outcome Rewards
- FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning
- Modeling Hierarchical Thinking in Large Reasoning Models
- PACR: Progressively Ascending Confidence Reward for LLM Reasoning
- Language Ranker: A Lightweight Ranking framework for LLM Decoding
- What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
- Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
- CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
- LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
- ExGRPO: Learning to Reason from Experience
- Cog-Rethinker: Hierarchical Metacognitive Reinforcement Learning for LLM Reasoning
- Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning
- Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
- Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
- Learning a Dense Reasoning Reward Model from Expert Demonstration via Inverse Reinforcement Learning
- Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
- Pinpointing crucial steps: Attribution-based Credit Assignment for Verifiable Reinforcement Learning
- Contrastive Weak-to-strong Generalization
- On the optimization dynamics of RLVR: Gradient gap and step size thresholds
- Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy Constraints
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
- Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
- TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning
- Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- Mitigating Forgetting Between Supervised and Reinforcement Learning Yields Stronger Reasoners
- TROLL: Trust Regions improve Reinforcement Learning for Large Language Models
- Graph-S3: Enhancing Agentic textual Graph Retrieval with Synthetic Stepwise Supervision
- Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
- Linking Process to Outcome: Conditional Reward Modeling for LLM Reasoning
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
- Diversity-Incentivized Exploration for Versatile Reasoning
- PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection
- Hybrid Reward Normalization for Process-supervised Non-verifiable Agentic Tasks
- Language Models Can Learn from Verbal Feedback Without Scalar Rewards
- Variational Reasoning for Language Models
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- Learning GUI Grounding with Spatial Reasoning from Visual Feedback
- Tree Search for LLM Agent Reinforcement Learning
- GRPO is Secretly a Process Reward Model
- Calibrated Reasoning: An Explanatory Verifier for Dynamic and Efficient Problem-Solving
- Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning
- Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models
- Agentic Reinforcement Learning with Implicit Step Rewards
- SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning
- THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
- BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
- Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
- Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning
- Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance
- SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control
- StepWiser: Stepwise Generative Judges for Wiser Reasoning
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models
- Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
- G2RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance
- Hard Examples Are All You Need: Maximizing GRPO Post-Training Under Annotation Budgets
- SSRL: Self-Search Reinforcement Learning
- Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization
- Reducing Cognitive Overhead in Tool Use via Multi-Small-Agent Reinforcement Learning
- EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning
- AMFT: Aligning LLM Reasoners by Meta-Learning the Optimal Imitation-Exploration Balance
- Sample-efficient LLM Optimization with Reset Replay
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
- TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
- Sotopia-RL: Reward Design for Social Intelligence
- Self-Questioning Language Models
- Enhancing Math Reasoning in Small-sized LLMs via Preview Difficulty-Aware Intervention
- RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
- VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
- EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
Related