vix.ing · top · new · best · stats

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents

2025/10/03 by Wonjoong Kim, Kim, Wonjoong, Yeonjun In +8 · 2 citations
Computer Science · #AI-based Problem Solving and Planning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Key (lock) #Multi-Agent Systems and Negotiation #Scalability #Simple (philosophy) #TRACE (psycholinguistics) #Topic Modeling #Trajectory

paper · pdf · doi:10.48550/arxiv.2510.02837

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2025/10/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

Although recent tool-augmented benchmarks involve complex requests, evaluation remains limited to answer matching, neglecting critical trajectory aspects like efficiency, hallucination, and adaptivity. The most straightforward method for evaluation is to compare an agent's trajectory with the ground-truth, but annotating all valid ground-truth trajectories is prohibitively expensive. In this manner, we introduce TRACE, a reference-free framework for the multi-dimensional evaluation of tool-augmented LLMs. By incorporating an evidence bank which accumulates knowledge from preceding steps, TRACE assesses an agent's reasoning trajectory effectively. To validate our framework, we develop a new meta-evaluation dataset with diverse and flawed trajectories, each labeled with multi-faceted performance scores. Our results confirm that TRACE accurately evaluates complex trajectories even with small open-source LLMs. Furthermore, we apply our method to evaluate the trajectories that agents produce while solving tool-augmented tasks, presenting previously unreported observations and their corresponding insights.

Citations

Cited by

Related