Evaluation and Benchmarking of LLM Agents: A Survey
2025/07/29 by Mohammadi, Mahmoud, Li, Yipeng, Lo, Jane +1 · 41 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
paper · doi:10.48550/arxiv.2507.21504
Abstract
The rise of LLM-based agents has opened new frontiers in AI applications, yet evaluating these agents remains a complex and underdeveloped area. This survey provides an in-depth overview of the emerging field of LLM agent evaluation, introducing a two-dimensional taxonomy that organizes existing work along (1) evaluation objectives -- what to evaluate, such as agent behavior, capabilities, reliability, and safety -- and (2) evaluation process -- how to evaluate, including interaction modes, datasets and benchmarks, metric computation methods, and tooling. In addition to taxonomy, we highlight enterprise-specific challenges, such as role-based access to data, the need for reliability guarantees, dynamic and long-horizon interactions, and compliance, which are often overlooked in current research. We also identify future research directions, including holistic, more realistic, and scalable evaluation. This work aims to bring clarity to the fragmented landscape of agent evaluation and provide a framework for systematic assessment, enabling researchers and practitioners to evaluate LLM agents for real-world deployment.
Citations
Cited by
- Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation
- Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents
- GeoDecider: An Evidence-Grounded Agent for Geological Interpretation via Deliberative Reasoning
- DREAM: Dynamic Red-teaming across Environments for AI Models
- Measuring Agents in Production
- AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets
- Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code Translation
- Failure Modes in LLM Systems: A System-Level Taxonomy for Reliable AI Applications
- How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
- Convergence dynamics of Agent-to-Agent Interactions with Misaligned objectives
- Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents
- AdvisingWise: Supporting Academic Advising in Higher Education Settings Through a Human-in-the-Loop Multi-Agent Framework
- Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks
- Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI
- Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- Embracing Trustworthy Brain-Agent Collaboration as Paradigm Extension for Intelligent Assistive Technologies
- Integrating Machine Learning into Belief-Desire-Intention Agents: Current Advances and Open Challenges
- The Trust Paradox in LLM-Based Multi-Agent Systems: When Collaboration Becomes a Security Vulnerability
- A Large-Language-Model Assisted Automated Scale Bar Detection and Extraction Framework for Scanning Electron Microscopic Images
- DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
- What Is Your Agent's GPA? A Framework for Evaluating Agent Goal-Plan-Action Alignment
- Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions
- Open Agent Specification (Agent Spec): A Unified Representation for AI Agents
- LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
- PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases
- Evaluating Agentic Bioinformatics through Function, Evidence, and Validation
- How Benchmarks Mis-Score Computer-Use Agents
- Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
- AI Compute Architecture and Evolution Trends
- CyberSleuth: Autonomous Blue-Team LLM Agent for Web Attack Forensics
- KonfAI: A Modular and Fully Configurable Framework for Deep Learning in Medical Imaging
- When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs
- NatureGAIA: Pushing the Frontiers of GUI Agents with a Challenging Benchmark and High-Quality Trajectory Dataset
- From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents
- A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?
- On the Wings of Imagination: Conflicting Script-based Multi-role Framework for Humor Caption Generation
- From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents
- AgentOS: From Application Silos to a Natural Language-Driven Data Ecosystem
- Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents
- ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
- EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents
Related