Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
2026/06/24 by Changdae Oh, Wendi Li, Seongheon Park +3 · 1 voice
Computer Science · #cs.LG #cs.AI
paper · pdf
Abstract
Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings remains prohibitively difficult: long-horizon interactions, irreversible actions, and stochastic environment feedback make both human annotation and Monte Carlo estimation infeasible at scale. In this work, we show that reinforcement learning (RL) post-training already provides the ingredients for effective step-level scoring, eliminating the need for dedicated reward model training altogether. Concretely, we derive an implicit advantage under a general stochastic Markov decision process, which we term progress advantage -- log-probability ratio between the RL-trained policy and its reference policy exactly recovers the optimal advantage function. This formulation makes the resulting signal annotation-free, domain-agnostic, and available as a byproduct of the standard RL post-training pipeline. We validate the effectiveness of the progress advantage across three different applications: test-time scaling, uncertainty quantification, and failure attribution on five benchmarks and four model families. Across all settings, it consistently outperforms confidence-based baselines and, despite requiring no task-specific training, surpasses dedicated trained reward models. We complement these results with deeper analyses on characteristics of progress advantage, offering practical guidance for adoption in real-world agentic systems.
Citations
- Qwen3.5-Omni Technical Report
- A Comedy of Estimators: On KL Regularization in RL Training of LLMs
- Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
- AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
- Reasoning with Sampling: Your Base Model is Smarter Than You Think
- Linking Process to Outcome: Conditional Reward Modeling for LLM Reasoning
- Agentic Reinforcement Learning with Implicit Step Rewards
- AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
- Deep Think with Confidence
- RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
- Spurious Rewards: Rethinking Training Signals in RLVR
- On a few pitfalls in KL divergence gradient estimation for RL
- τ2-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning
- Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening
- Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision Making
- Qwen3 Technical Report
- Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
- Process Reward Models That Think
- Information-Theoretic Reward Decomposition for Generalizable RLHF
- Understanding R1-Zero-Like Training: A Critical Perspective
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Scalable Best-of-N Selection for Large Language Models via Self-Certainty
- Process Reward Models for LLM Agents: Practical Framework and Directions
- VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data
- Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
- Process Reinforcement through Implicit Rewards
- The Lessons of Developing Process Reward Models in Mathematical Reasoning
- Qwen2.5 Technical Report
- The Open Source Advantage in Large Language Models (LLMs)
- ProcessBench: Identifying Process Errors in Mathematical Reasoning
- Self-Improvement in Language Models: The Sharpening Mechanism
- Free Process Rewards without Process Labels
- Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks
- Process Reward Model with Q-Value Rankings
- Agent-as-a-Judge: Evaluate Agents with Agents
- DaWin: Training-free Dynamic Weight Interpolation for Robust Adaptation
- LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
- Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
- AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
- The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
- Robust Adaptation of Foundation Models with Black-Box Visual Prompting
- Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
- Direct Multi-Turn Preference Optimization for Language Agents
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Improve Mathematical Reasoning in Language Models by Automated Process Supervision
- Adaptive In-conversation Team Building for Language Model Agents
- SimPO: Simple Preference Optimization with a Reference-Free Reward
- RLHF Workflow: From Reward Modeling to Online RLHF
- From r to Q^*: Your Language Model is Secretly a Q-Function
- Model Stock: All we need is just a few fine-tuned models
- Evolutionary optimization of model merging recipes
- ORPO: Monolithic Preference Optimization without Reference Model
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- Self-Rewarding Language Models
- Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus
- GAIA: a benchmark for General AI Assistants
- OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning
- Black-Box Prompt Optimization: Aligning Large Language Models without Model Training
- Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
- Self-Guard: Empower the LLM to Safeguard Itself
- A General Theoretical Paradigm to Understand Learning from Human Preferences
- AdaMerging: Adaptive Model Merging for Multi-Task Learning
- Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
- Contrastive Decoding Improves Reasoning in Large Language Models
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
- Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- TIES-Merging: Resolving Interference When Merging Models
- Let's Verify Step by Step
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Inverse Preference Learning: Preference-based RL without a Reward Function
- Can Large Language Models Be an Alternative to Human Evaluations?
- Neglected Free Lunch -- Learning Image Classifiers Using Annotation Byproducts
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Affective Coherence Monitoring for Transformer-Based Language Models
- Editing Models with Task Arithmetic
- Solving math word problems with process- and outcome-based feedback
- Contrastive Decoding: Open-ended Text Generation as Optimization
- Scaling Laws for Reward Model Overoptimization
- Patching open-vocabulary models by interpolating weights
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
- Learning Fair Representation via Distributional Contrastive Disentanglement
- RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- STaR: Bootstrapping Reasoning With Reasoning
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
- Black-Box Tuning for Language-Model-as-a-Service
- Training Verifiers to Solve Math Word Problems
- Robust fine-tuning of zero-shot models
- Rank consistent ordinal regression for neural networks with application to age estimation
- Noise Contrastive Estimation and Negative Sampling for Conditional Models: Consistency and Statistical Efficiency
- Proximal Policy Optimization Algorithms
- Snapshot Ensembles: Train 1, get M for free
- Reinforcement Learning with Deep Energy-Based Policies
- Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles
- Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning
- Training Products of Experts by Minimizing Contrastive Divergence
- GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning
- Reinforcement Learning: An Introduction
Discussions
Related