The Lessons of Developing Process Reward Models in Mathematical Reasoning
2025/01/13 by Zhenru Zhang, Zhang, Zhenru, Chujie Zheng +15 · 2 voices · 208 citations
Computer Science · #cs.CL #cs.AI #cs.LG
paper · pdf · doi:10.48550/arxiv.2501.07301
Abstract
Process Reward Models (PRMs) emerge as a promising approach for process supervision in mathematical reasoning of Large Language Models (LLMs), which aim to identify and mitigate intermediate errors in the reasoning processes. However, the development of effective PRMs faces significant challenges, particularly in data annotation and evaluation methodologies. In this paper, through extensive experiments, we demonstrate that commonly used Monte Carlo (MC) estimation-based data synthesis for PRMs typically yields inferior performance and generalization compared to LLM-as-a-judge and human annotation methods. MC estimation relies on completion models to evaluate current-step correctness, leading to inaccurate step verification. Furthermore, we identify potential biases in conventional Best-of-N (BoN) evaluation strategies for PRMs: (1) The unreliable policy models generate responses with correct answers but flawed processes, leading to a misalignment between the evaluation criteria of BoN and the PRM objectives of process verification. (2) The tolerance of PRMs of such responses leads to inflated BoN scores. (3) Existing PRMs have a significant proportion of minimum scores concentrated on the final answer steps, revealing the shift from process to outcome-based assessment in BoN Optimized PRMs. To address these challenges, we develop a consensus filtering mechanism that effectively integrates MC estimation with LLM-as-a-judge and advocates a more comprehensive evaluation framework that combines response-level and step-level metrics. Based on the mechanisms, we significantly improve both model performance and data efficiency in the BoN evaluation and the step-wise error identification task. Finally, we release a new state-of-the-art PRM that outperforms existing open-source alternatives and provides practical guidelines for future research in building process supervision models.
Cited by
- Verify, Repair, Repeat, or Stop? Robust Stopping for Noisy Verify-Repair Loops in LLM Agents
- UnMaskFork: Test-Time Scaling for Masked Diffusion via Deterministic Action Branching
- LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning
- Agent Step Value: Auditing Evaluator-Channel Reversals in Black-Box Agent Traces
- Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models
- Scaling Evaluation-time Compute with Reasoning Models as Evaluators
- Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- Inference-Time Scaling for Generalist Reward Modeling
- Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits
- Training LLMs with LogicReward for Faithful and Rigorous Reasoning
- Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
- Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
- Stepwise Think-Critique: A Unified Framework for Robust and Interpretable LLM Reasoning
- Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning
- OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification
- PyFi: Toward Pyramid-like Financial Image Understanding for VLMs via Adversarial Agents
- Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
- SPARK: Stepwise Process-Aware Rewards for Reference-Free Reinforcement Learning
- SR-GRPO: Stable Rank as an Intrinsic Geometric Reward for Large Language Model Alignment
- DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action
- Video Generation Models Are Good Latent Reward Models
- HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs
- In-Token Rationality Optimization: Towards Accurate and Concise LLM Reasoning via Self-Feedback
- DeCoRL: Decoupling Reasoning Chains via Parallel Sub-Step Generation and Cascaded Reinforcement for Interpretable and Scalable RLHF
- ReProbe: Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
- Lookahead Unmasking Elicits Accurate Decoding in Diffusion Language Models
- VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought Supervision
- Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?
- MAS-ProVe: Understanding the Process Verification of Multi-Agent Systems
- TrajSelector: Harnessing Latent Representations for Efficient and Effective Best-of-N in Large Reasoning Model
- FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning
- Sample By Step, Optimize By Chunk: Chunk-Level GRPO For Text-to-Image Generation
- Exploring Generative Process Reward Modeling for Semi-Structured Data: A Case Study of Table Question Answering
- Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
- SmartSwitch: Advancing LLM Reasoning by Overcoming Underthinking via Promoting Deeper Thought Exploration
- No Compute Left Behind: Rethinking Reasoning and Sampling with Masked Diffusion Models
- Adaptive Coopetition: Leveraging Coarse Verifier Signals for Resilient Multi-Agent LLM Reasoning
- Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
- DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMs
- A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning
- CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
- LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
- GroundedPRM: Tree-Guided and Fidelity-Aware Process Reward Modeling for Step-Level Reasoning
- Qwen3Guard Technical Report
- Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- Optimal Aggregation of LLM and PRM Signals for Efficient Test-Time Scaling
- BoN Appetit Team at LeWiDi-2025: Best-of-N Test-time Scaling Can Not Stomach Annotation Disagreements (Yet)
- Towards Inference-time Scaling for Continuous Space Reasoning
- Enhancing LLM Reasoning via Non-Human-Like Reasoning Path Preference Optimization
- Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning
- Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
- ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering
- On the Provable Performance Guarantee of Efficient Reasoning Models
- CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
- A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
- Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
- Hybrid Reinforcement: When Reward Is Sparse, It's Better to Be Dense
- Step-Aware Policy Optimization for Reasoning in Diffusion Large Language Models
- TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning
- Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Multi-Stream Environments
- Beyond Token Length: Step Pruner for Efficient and Accurate Reasoning in Large Language Models
- Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
- Rethinking Reward Models for Multi-Domain Test-Time Scaling
- Linking Process to Outcome: Conditional Reward Modeling for LLM Reasoning
- SCUBA: Salesforce Computer Use Benchmark
- Hybrid Reward Normalization for Process-supervised Non-verifiable Agentic Tasks
- From Faithfulness to Correctness: Generative Reward Models that Think Critically
- Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
- ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling
- Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned
- Dynamic Experts Search: Enhancing Reasoning in Mixture-of-Experts LLMs at Test Time
- Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts
- GRPO is Secretly a Process Reward Model
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
- SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning
- Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning
- Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents
- THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
- LATTS: Locally Adaptive Test-Time Scaling
- All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model Reasoning
- ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
- RAFFLES: Reasoning-based Attribution of Faults for LLM Systems
- Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
- GradeSQL: Test-Time Inference with Outcome Reward Models for Text-to-SQL Generation from Large Language Models
- ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding
- StepWiser: Stepwise Generative Judges for Wiser Reasoning
- Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models
- Trust but Verify! A Survey on Verification Design for Test-time Scaling
- Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
- Improving Value-based Process Verifier via Low-Cost Variance Reduction
- Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
- Reducing Cognitive Overhead in Tool Use via Multi-Small-Agent Reinforcement Learning
- From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR
- AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models
- TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
- Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction
- CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
- VRPRM: Process Reward Modeling via Visual Reasoning
- CTTS: Collective Test-Time Scaling
- PentestJudge: Judging Agent Behavior Against Operational Requirements
- CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment
- Test-time Prompt Intervention
- Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical Reasoning
- The Bidirectional Process Reward Model
- Good Learners Think Their Thinking: Generative PRM Makes Large Reasoning Model More Efficient Math Learner
- ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge
- Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models
- Post-Completion Learning for Language Models
- CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
- RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback
- A Primer in Post-Training Reasoning Data: What We Know About How It Works
- GenSelect: A Generative Approach to Best-of-N
- Probabilistic Soundness Guarantees in LLM Reasoning Chains
- Med-REFL: Medical Reasoning Enhancement via Self-Corrected Fine-grained Reflection
- Learning to Reason Across Parallel Samples for LLM Reasoning
- Large Language Models Have Intrinsic Meta-Cognition, but Need a Good Lens
- OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique
- Inference-Time Scaling of Diffusion Language Models with Particle Gibbs Sampling
- Bradley-Terry and Multi-Objective Reward Modeling Are Complementary
- Review, Remask, Refine (R3): Process-Guided Block Diffusion for Text Generation
- PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
- Test-Time Scaling with Reflective Generative Model
- Improving Rationality in the Reasoning Process of Language Models through Self-playing Game
- OptScale: Probabilistic Optimality for Inference-time Scaling
- Lost at the Beginning of Reasoning
- Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
- ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMs
- Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search
- Strategic Scaling of Test-Time Compute: A Bandit Learning Approach
SPECS: Faster Test-Time Scaling through Speculative Drafts- Graph of Verification: Structured Verification of LLM Reasoning with Directed Acyclic Graphs
- TongSearch-QR: Reinforced Query Reasoning for Retrieval
- PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier
- Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
- Saffron-1: Safety Inference Scaling
- Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers
- FreePRM: Training Process Reward Models Without Ground Truth Process Labels
- The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
- How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning
- Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
- MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
- Socratic-PRMBench: Benchmarking Process Reward Models with Systematic Reasoning Patterns
- AutoL2S: Auto Long-Short Reasoning for Efficient Large Language Models
- What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoning
- LASER: Stratified Selective Sampling for Instruction Tuning with Dedicated Scoring Strategy
- Step-Wise Formal Verification for LLM-Based Mathematical Problem Solving
- Beyond Templates: Dynamic Adaptation of Reasoning Demonstrations via Feasibility-Aware Exploration
- Can Past Experience Accelerate LLM Reasoning?
- Enhancing Logical Reasoning in Language Models via Symbolically-Guided Monte Carlo Process Supervision
- Temporal Sampling for Forgotten Reasoning in LLMs
- Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision
- Training-Free Multi-Step Audio Source Separation
- Interleaved Reasoning for Large Language Models via Reinforcement Learning
- Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
- VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
- Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence
- ProgRM: Build Better GUI Agents with Progress Rewards
- Reward Model Generalization for Compute-Aware Test-Time Reasoning
- Value-Guided Search for Efficient Chain-of-Thought Reasoning
- p2-TQA: A Process-based Preference Learning Framework for Self-Improving Table Question Answering Models
- SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward
- Generalizable Process Reward Models via Formally Verified Training Data
- Swarm Intelligence Enhanced Reasoning: A Density-Driven Framework for LLM-Based Multi-Agent Optimization
- Unearthing Gems from Stones: Policy Optimization with Negative Sample Augmentation for LLM Reasoning
- RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning
- Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning
- Reward Reasoning Model
- A*-Decoding: Token-Efficient Inference Scaling
- Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
- Walking the Tightrope: Disentangling Beneficial and Detrimental Drifts in Non-Stationary Custom-Tuning
- Fractured Chain-of-Thought Reasoning
- MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision
- Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective
- Thinking Short and Right Over Thinking Long: Serving LLM Reasoning Efficiently and Accurately
- Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Reward Design
- Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
- Real-Time Verification of Embodied Reasoning for Generative Skill Acquisition
- Dist2ill: Distributional Distillation for One-Pass Uncertainty Estimation in Large Language Models
- Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
- Fusing Bidirectional Chains of Thought and Reward Mechanisms A Method for Enhancing Question-Answering Capabilities of Large Language Models for Chinese Intangible Cultural Heritage
- Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation
- Reinforcement Learning without Ground-Truth Solutions can Improve LLMs
- An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- Pencil Puzzle Bench: A Benchmark for Multi-Step Verifiable Reasoning
- DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards
- Online Safety Monitoring for LLMs
- Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
- Scalable Token-Level Hallucination Detection in Large Language Models
- MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild
- Learning to Draw ASCII Improves Spatial Reasoning in Language Models
- Process Reward Agents for Steering Knowledge-Intensive Reasoning
- CausalOPD: First-Wrong-Step Supervision for Distilling Causal Chain Reasoning
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
- T1: Tool-integrated Verification for Test-time Compute Scaling in Small Language Models
- Tree of Thoughts as a Classical Heuristic Search Problem: Formal Foundations and Design Patterns
- Process Reward Models That Think
- Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning
- Large models for machinery fault diagnosis: Current advances and future directions
- Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning
- Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
- Efficient MAP Estimation of LLM Judgment Performance with Prior Transfer
- Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning
- Correct Answers from Sound Reasoning: Verifiable Process Supervision for Language Models
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Efficient Process Reward Model Training via Active Learning
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
Discussions
Related