On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
2025/12/04 by Yu, Yue, Di, Qiwei, Gu, Quanquan +1
#FOS: Computer and information sciences #Machine Learning (cs.LG)
paper · doi:10.48550/arxiv.2512.04558
Abstract
Test-time compute (TTC) has become an increasingly prominent paradigm for enhancing large language models (LLMs). Despite the empirical success of methods such as best-of-n (BoN) sampling and sequential revision, their fundamental limits remain unclear. We address this gap by analyzing a mixture-of-reference policy model and proving that standard BoN is inherently suboptimal. To move closer to the optimal frontier, we study reward-filtered sequential inference, a simple procedure that selectively incorporates only high-reward generations into the context. This mechanism concentrates computation on superior policy candidates and suppresses inferior ones. On the theoretical side, we show that reward-filtered sequential inference yields strictly stronger guarantees than standard TTC paradigms. On the empirical side, we evaluate such an inference strategy across diverse benchmarks and observe consistent improvements over widely used approaches, demonstrating the practical effectiveness of our framework.
Citations
- Maximizing Prefix-Confidence at Test-Time Efficiently Improves Mathematical Reasoning
- Strategic Scaling of Test-Time Compute: A Bandit Learning Approach
- Provably Learning from Language Feedback
- PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier
- Sample Complexity and Representation Ability of Test-time Scaling Paradigms
- OpenThoughts: Data Recipes for Reasoning Models
- First Finish Search: Efficient Test-Time Scaling in Large Language Models
- Process Reward Models That Think
- M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
- Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
- Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration
- Self-Training Elicits Concise Reasoning in Large Language Models
- Scalable Best-of-N Selection for Large Language Models via Self-Certainty
- Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning under Misspecification
- Uncertainty-Aware Step-wise Verification with Generative Reward Models
- Evolving Deeper LLM Thinking
- AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
- Qwen2.5 Technical Report
- Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
- The Limits of Inference Scaling Through Resampling
- How to Evaluate Reward Models for RLHF
- Guaranteed Generation from Large Language Models
- Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Is Behavior Cloning All You Need? Understanding Horizon in Imitation Learning
- BOND: Aligning LLMs with Best-of-N Distillation
- The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
- Transformers Can Learn Temporal Difference Methods for In-Context Reinforcement Learning
- Small Language Models Need Strong Verifiers to Self-Correct Reasoning
- Theoretical guarantees on the best-of-n alignment policy
- Universal Self-Consistency for Large Language Model Generation
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models
- Supervised Pretraining Can Learn In-Context Reinforcement Learning
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm Selection
- Let's Verify Step by Step
- What and How does In-Context Learning Learn? Bayesian Model Averaging, Parameterization, and Generalization
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Teaching Large Language Models to Self-Debug
- Self-Refine: Iterative Refinement with Self-Feedback
- Reflexion: Language Agents with Verbal Reinforcement Learning
- The Learnability of In-Context Learning
- Rewarding Chatbots for Real-World Engagement with Millions of Users
- Transformers as Algorithms: Generalization and Stability in In-context Learning
- Discovering Language Model Behaviors with Model-Written Evaluations
- Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
- Scaling Laws for Reward Model Overoptimization
- ReAct: Synergizing Reasoning and Acting in Language Models
- Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
- Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- WebGPT: Browser-assisted question-answering with human feedback
- An Explanation of In-context Learning as Implicit Bayesian Inference
- Measuring Mathematical Problem Solving With the MATH Dataset
- Learning to summarize from human feedback
- Reward Tampering Problems and Solutions in Reinforcement Learning: A\n Causal Influence Diagram Perspective
- Concrete Problems in AI Safety
- Bayesian model averaging: a tutorial (with comments by M. Clyde, David Draper and E. I. George, and a rejoinder by the authors
- Adaptive Mixtures of Local Experts
- Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
Related