CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
2025/11/04 by Liu, Jiayu, Qian, Cheng, Su, Zhaochen +4
Computer Science · Medicine · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
paper · doi:10.48550/arxiv.2511.02734
openalex publication_date 2025/11/04 · openalex created_date 2025/11/06 · openalex updated_date 2026/07/28
Abstract
Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and adaptability. This neglects a crucial capability: agents' ability to devise and adjust cost-optimal plans in response to changing environments. To bridge this gap, we introduce CostBench, a scalable, cost-centric benchmark designed to evaluate agents' economic reasoning and replanning abilities. Situated in the travel-planning domain, CostBench comprises tasks solvable via multiple sequences of atomic and composite tools with diverse, customizable costs. It also supports four types of dynamic blocking events, such as tool failures and cost changes, to simulate real-world unpredictability and necessitate agents to adapt in real time. Evaluating leading open-sourced and proprietary models on CostBench reveals a substantial gap in cost-aware planning: agents frequently fail to identify cost-optimal solutions in static settings, with even GPT-5 achieving less than 75% exact match rate on the hardest tasks, and performance further dropping by around 40% under dynamic conditions. By diagnosing these weaknesses, CostBench lays the groundwork for developing future agents that are both economically rational and robust.
Citations
- Budget-Aware Tool-Use Enables Effective Agent Scaling
- Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
- WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
- ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding
- BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens
- R-ConstraintBench: Evaluating LLMs on NP-Complete Scheduling
- WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent
- UserBench: An Interactive Gym Environment for User-Centric Agents
- Diversity-Enhanced Reasoning for Subjective Questions
- Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
- Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
- Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models
- Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
- Revisiting Epistemic Markers in Confidence Estimation: Can Markers Accurately Reflect Large Language Models' Uncertainty?
- WebDancer: Towards Autonomous Information Seeking Agency
- AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware Budgeting
- OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
- Acting Less is Reasoning More! Teaching Model to Act Efficiently
- ToolRL: Reward is All Tool Learning Needs
- ACPBench Hard: Unrestrained Reasoning about Action, Change, and Planning
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Advancing Language Multi-Agent Learning with Credit Re-Assignment for Interactive Environment Generalization
- SMART: Self-Aware Agent for Tool Overuse Mitigation
- MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
- ACEBench: Who Wins the Match Point in Tool Usage?
- ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
- CATP-LLM: Empowering Large Language Models for Cost-Aware Tool Planning
- ACPBench: Reasoning about Action, Change, and Planning
- BudgetMLAgent: A Cost-Effective LLM Multi-Agent system for Automating Machine Learning Tasks
- ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning Strategies
- Bootstrap confidence intervals: A comparative simulation study
- TravelPlanner: A Benchmark for Real-World Planning with Language Agents
- Planning, Creation, Usage: Benchmarking LLMs for Comprehensive Tool Utilization in Real-World Complex Scenarios
- Put Your Money Where Your Mouth Is: Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena
- EcoAssistant: Using LLM Assistant More Affordably and Accurately
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- SayCanPay: Heuristic Planning with Large Language Models using Learnable Domain Knowledge
- Counterfactually Auditable Lifecycle Certification for Autonomous Agents
- ToolQA: A Dataset for LLM Question Answering with External Tools
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- LLaMA: Open and Efficient Foundation Language Models
- Toolformer: Language Models Can Teach Themselves to Use Tools
- PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change
- Edit Distance Cannot Be Computed in Strongly Subquadratic Time (unless\n SETH is false)
- Levenshtein Distance Technique in Dictionary Lookup Methods: An Improved Approach
Cited by
Related