Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
2024/08/01 by Yangzhen Wu, Zhiqing Sun, Wu, Yangzhen +7 · 1 voice · 195 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Natural Language Processing Techniques #cs.AI
paper · pdf · doi:10.48550/arxiv.2408.00724
openalex publication_date 2024/08/01 · arxiv published 2024/08/01 · openalex created_date 2024/08/04 · arxiv updated 2025/03/03 · openalex updated_date 2026/07/28
Abstract
While the scaling laws of large language models (LLMs) training have been extensively studied, optimal inference configurations of LLMs remain underexplored. We study inference scaling laws (aka test-time scaling laws) and compute-optimal inference, focusing on the trade-offs between model sizes and generating additional tokens with different inference strategies. As a first step towards understanding and designing compute-optimal inference methods, we studied cost-performance trade-offs for inference strategies such as greedy search, majority voting, best-of-n, weighted voting, and two different tree search algorithms, using different model sizes and compute budgets. Our findings suggest that scaling inference compute with inference strategies can be more computationally efficient than scaling model parameters. Additionally, smaller models combined with advanced inference algorithms offer Pareto-optimal trade-offs in cost and performance. For example, the Llemma-7B model, when paired with our novel tree search algorithm, consistently outperforms the Llemma-34B model across all tested inference strategies on the MATH benchmark. We hope these insights contribute to a deeper understanding of inference scaling laws (test-time scaling laws) for LLMs.
Cited by
- Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling
- GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
- Meta-RL Induces Exploration in Language Agents
- Scaling Laws for Energy Efficiency of Local LLMs
- Reliable agent engineering should integrate machine-compatible organizational principles
- On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
- Think in Parallel, Answer as One: Logit Averaging for Open-Ended Reasoning
- Beware of Reasoning Overconfidence: Pitfalls in the Reasoning Process for Multi-solution Tasks
- Zero-Overhead Introspection for Adaptive Test-Time Compute
- ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
- Budget-Aware Tool-Use Enables Effective Agent Scaling
- Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
- TRIM: Scalable 3D Gaussian Diffusion Inference with Temporal and Spatial Trimming
- Incorporating Self-Rewriting into Large Language Model Reasoning Reinforcement
- Extending Test-Time Scaling: A 3D Perspective with Context, Batch, and Turn
- Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models
- Efficient Thought Space Exploration Through Strategic Intervention
- DeCoRL: Decoupling Reasoning Chains via Parallel Sub-Step Generation and Cascaded Reinforcement for Interpretable and Scalable RLHF
- Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
- CGES: Confidence-Guided Early Stopping for Efficient and Accurate Self-Consistency
- Test-time Scaling of LLMs: A Survey from A Subproblem Structure Perspective
- e1: Learning Adaptive Control of Reasoning Effort
- Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph
- Evaluating the Role of Verifiers in Test-Time Scaling for Legal Reasoning Tasks
- Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
- FrugalPrompt: Reducing Contextual Overhead in Large Language Models via Token Attribution
- Boosting Accuracy and Efficiency of Budget Forcing in LLMs via Reinforcement Learning for Mathematical Reasoning
- Language Ranker: A Lightweight Ranking framework for LLM Decoding
- EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
- 3D Optimization for AI Inference Scaling: Balancing Accuracy, Cost, and Latency
- MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
- Dual-Weighted Reinforcement Learning for Generative Preference Modeling
- ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling
- Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models
- Improving Text-to-Image Generation with Input-Side Inference-Time Scaling
- EAGer: Entropy-Aware GEneRation for Adaptive Inference-Time Scaling
- Cog-Rethinker: Hierarchical Metacognitive Reinforcement Learning for LLM Reasoning
- Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning
- Mitigating Overthinking through Reasoning Shaping
- Upfront Chain-of-Thought: A Cooperative Framework for Chain-of-Thought Compression
- SurveyG: A Multi-Agent LLM Framework with Hierarchical Citation Graph for Automated Survey Generation
- Tool-Augmented Policy Optimization: Synergizing Reasoning and Adaptive Tool Use with Reinforcement Learning
- Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
- On the Role of Temperature Sampling in Test-Time Scaling
- ARISE: An Adaptive Resolution-Aware Metric for Test-Time Scaling Evaluation in Large Reasoning Models
- PatternKV: Flattening KV Representation Expands Quantization Headroom
- GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
- HALO: Memory-Centric Heterogeneous Accelerator with 2.5D Integration for Low-Batch LLM Inference
- Generalized Parallel Scaling with Interdependent Generations
- Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining
- Entropy After ⟨
/Think ⟩ for reasoning model early exiting - Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training
- DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search
- Adaptive Test-Time Reasoning via Reward-Guided Dual-Phase Search
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling
- Evaluating the Robustness of Chinchilla Compute-Optimal Scaling
- Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm
- Training Optimal Large Diffusion Language Models
- Pretraining Scaling Laws for Generative Evaluations of Language Models
- Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
- From Harm to Help: Turning Reasoning In-Context Demos into Assets for Reasoning LMs
- Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM
- Transformers Can Learn Connectivity in Some Graphs but Not Others
- Best-of-∞ -- Asymptotic Performance of Test-Time LLM Ensembling
- Parallel Thinking, Sequential Answering: Bridging NAR and AR for Efficient Reasoning
- Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs
- Investigating Test-Time Scaling with Reranking for Machine Translation
- TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
- ATTS: Asynchronous Test-Time Scaling via Conformal Prediction
- FLAME: A Serving System Optimized for Large-Scale Generative Recommendation with Efficiency
- Virtual Agent Economies
- Uncovering Scaling Laws for Large Language Models via Inverse Problems
- Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
- Outcome-based Exploration for LLM Reasoning
- The Majority is not always right: RL training for solution aggregation
- CoT-Space: A Theoretical Framework for Internal Slow-Thinking via Reinforcement Learning
- GRAM-R2: Self-Training Generative Foundation Reward Models for Reward Reasoning
- EviNote-RAG: Enhancing RAG Models via Answer-Supportive Evidence Notes
- LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
- Democratizing Agentic AI with Fast Test-Time Scaling on the Edge
- PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
- Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
- SLIM: Subtrajectory-Level Elimination for More Effective Reasoning
- SynthCoder: A Synthetical Strategy to Tune LLMs for Code Completion
- Deep Think with Confidence
- Long Chain-of-Thought Reasoning Across Languages
- Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
- Navigating the Exploration-Exploitation Tradeoff in Inference-Time Scaling of Diffusion Models
- Train Long, Think Short: Curriculum Learning for Efficient Reasoning
- TeamMedAgents: Pareto-Efficient Multi-Agent Medical Reasoning Through Teamwork Theory
- Uncertainty-Aware Semantic Decoding for LLM-Based Sequential Recommendation
- Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction
- Reasoning as a Resource: Optimizing Fast and Slow Thinking in Code Generation Models
- Lessons from complex systems science for AI governance
- Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models
- AgentTTS: Large Language Model Agent for Test-time Compute-optimal Scaling Strategy in Complex Tasks
- CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
- Computational Arbitrage in AI Model Markets
- V1: Unifying Generation and Self-Verification for Parallel Reasoners
- AlgoVeri: An Aligned Benchmark for Verified Code Generation on Classical Algorithms
- GenSelect: A Generative Approach to Best-of-N
- The Serial Scaling Hypothesis
- A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search
- A Survey on Large Language Models for Mathematical Reasoning
- DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models
- AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model
- PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem Solving
- Why is Your Language Model a Poor Implicit Reward Model?
- Agentic-R1: Distilled Dual-Strategy Reasoning
- APPO: Agentic Procedural Policy Optimization
- Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
- Reasoning as an Adaptive Defense for Safety
- InvisibleInk: High-Utility and Low-Cost Text Generation with Differential Privacy
- Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language Models
- Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
- A Survey of LLM Inference Systems
- CoMind: Towards Community-Driven Agents for Machine Learning Engineering
- Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs
- Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
- Maximal Update Parametrization and Zero-Shot Hyperparameter Transfer for Fourier Neural Operators
- ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMs
- From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
- GRAM: A Generative Foundation Reward Model for Reward Generalization
- MLDebugging: Towards Benchmarking Code Debugging Across Multi-Library Scenarios
- Scaling Test-time Compute for LLM Agents
- TongSearch-QR: Reinforced Query Reasoning for Retrieval
- ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization
- Predictable Scale: Part II, Farseer: A Refined Scaling Law in Large Language Models
- Route-and-Reason: Scaling Large Language Model Reasoning with Reinforced Model Router
- Scalable Chain of Thoughts via Elastic Reasoning
- Saffron-1: Safety Inference Scaling
- Sample Complexity and Representation Ability of Test-time Scaling Paradigms
- Kinetics: Rethinking Test-Time Scaling Laws
- Crosslingual Reasoning through Test-Time Scaling
- Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
- Dual-Process Image Generation
- Large language models can learn and generalize steganographic chain-of-thought under process supervision
- The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
- ProtInvTree: Deliberate Protein Inverse Folding with Reward-guided Tree Search
- Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
- Inference-time Scaling of Diffusion Models through Classical Search
- Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization
- Can Past Experience Accelerate LLM Reasoning?
- Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
- A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law
- Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models
- Temporal Sampling for Forgotten Reasoning in LLMs
- To CoT or To Loop? A Formal Comparison Between Chain-of-Thought and Looped Transformers
- Partition Generative Modeling: Masked Modeling Without Masks
- Inference Compute-Optimal Video Vision Language Models
- First Finish Search: Efficient Test-Time Scaling in Large Language Models
- Reward Model Generalization for Compute-Aware Test-Time Reasoning
- FlashForge: Ultra-Efficient Prefix-Aware Attention for LLM Decoding
- Scaling Image and Video Generation via Test-Time Evolutionary Search
- T2: An Adaptive Test-Time Scaling Strategy for Contextual Question Answering
- TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling
- Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language Models
- Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
- Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
- Reward Reasoning Model
- Efficient Agent Training for Computer Use
- MR. Judge: Multimodal Reasoner as a Judge
- Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative Verifier
- J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge
- Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
- Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs
- Follow the Path: Reasoning over Knowledge Graph Paths to Improve Large Language Model Factuality
- HAPO: Training Language Models to Reason Concisely via History-Aware Policy Optimization
- Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
- HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation
- Parallel Scaling Law for Language Models
- Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
- CodePDE: An Inference Framework for LLM-driven PDE Solver Generation
- Scalable LLM Math Reasoning Acceleration with Low-rank Distillation
- Position: Enough of Scaling LLMs! Lets Focus on Downscaling
- Architecting Trust in Artificial Epistemic Agents
- When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling
- BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution
- Self-Improving Language Models with Bidirectional Evolutionary Search
- When Less is Enough: Efficient Inference via Collaborative Reasoning
- Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
- ThinkFL: Self-Refining Failure Localization for Microservice Systems via Reinforcement Fine-Tuning
- DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training
- Process Reward Models That Think
- Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
- Efficient Reasoning Models: A Survey
- Weight Ensembling Improves Reasoning in Language Models
- Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning
- Enhancing Mathematical Reasoning in Large Language Models with Self-Consistency-Based Hallucination Detection
- Leveraging Reasoning Model Answers to Enhance Non-Reasoning Model Capability
- ThoughtProbe: Classifier-Guided Thought Space Exploration Leveraging LLM Intrinsic Reasoning
- Sample, Don't Search: Rethinking Test-Time Alignment for Language Models
Discussions
Related