Survey on Evaluation of LLM-based Agents
2025/03/20 by Asaf Yehudai, Yehudai, Asaf, Lilach Eden +13 · 2 voices · 45 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2503.16416
arxiv published 2025/03/20 · arxiv updated 2026/04/23
Abstract
LLM-based agents represent a paradigm shift in AI, enabling autonomous systems to plan, reason, and use tools while interacting with dynamic environments. This paper provides the first comprehensive survey of evaluation methods for these increasingly capable agents. We analyze the field of agent evaluation across five perspectives: (1) Core LLM capabilities needed for agentic workflows, like planning, and tool use; (2) Application-specific benchmarks such as web and SWE agents; (3) Evaluation of generalist agents; (4) Analysis of agent benchmarks' core dimensions; and (5) Evaluation frameworks and tools for agent developers. Our analysis reveals current trends, including a shift toward more realistic, challenging evaluations with continuously updated benchmarks. We also identify critical gaps that future research must address, particularly in assessing cost-efficiency, safety, and robustness, and in developing fine-grained, scalable evaluation methods.
Cited by
- Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation
- Towards Responsible and Explainable AI Agents with Consensus-Driven Reasoning
- A Practical Guide for Designing, Developing, and Deploying Production-Grade Agentic AI Workflows
- Towards a Science of Scaling Agent Systems
- Stochasticity in Agentic Evaluations: Quantifying Inconsistency with Intraclass Correlation
- Measuring Agents in Production
- An Empirical Study of Agent Developer Practices in AI Agent Frameworks
- AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets
- Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
- Rethinking the Reliability of Multi-agent System: A Perspective from Byzantine Fault Tolerance
- Environment Scaling for Interactive Agentic Experience Collection: A Survey
- Standardization of Psychiatric Diagnoses -- Role of Fine-tuned LLM Consortium and OpenAI-gpt-oss Reasoning LLM Enabled Decision Support System
- Agentic AI: A Comprehensive Survey of Architectures, Applications, and Future Directions
- EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning
- Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents
- ColorEcosystem: Powering Personalized, Standardized, and Trustworthy Agentic Service in massive-agent Ecosystem
- Integrating Machine Learning into Belief-Desire-Intention Agents: Current Advances and Open Challenges
- ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering
- Semantic Intelligence: A Bio-Inspired Cognitive Framework for Embodied Agents
- To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
- Static Sandboxes Are Inadequate: Modeling Societal Complexity Requires Open-Ended Co-Evolution in LLM-Based Multi-Agent Simulations
- DeepResearchGuard: Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety
- Past, Present, and Future of Bug Tracking in the Generative AI Era
- Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions
- BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
- MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
- Evaluating and Mitigating Errors in LLM-Generated Web API Integrations
- How Benchmarks Mis-Score Computer-Use Agents
- LLM Agents at the Roundtable: A Multi-Perspective and Dialectical Reasoning Framework for Essay Scoring
- AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- AgenticIE: An Adaptive Agent for Information Extraction from Complex Regulatory Documents
- Throttling Web Agents Using Reasoning Gates
- AI Compute Architecture and Evolution Trends
- MindGuard: Intrinsic Decision Inspection for Securing LLM Agents Against Metadata Poisoning
- Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol
- AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios
- S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner
- Collab-REC: An LLM-based Agentic Framework for Balancing Recommendations in Tourism
- HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds
- Standardization of Neuromuscular Reflex Analysis -- Role of Fine-Tuned Vision-Language Model Consortium and OpenAI gpt-oss Reasoning LLM Enabled Decision Support System
- PentestJudge: Judging Agent Behavior Against Operational Requirements
- Pro2Guard: Proactive Runtime Enforcement of LLM Agent Safety via Probabilistic Model Checking
- MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them
- Agent0: Leveraging LLM Agents to Discover Multi-value Features from Text for Enhanced Recommendations
Discussions
Related