AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
2024/01/24 by Chang Ma, Junlei Zhang, Ma, Chang +15 · 91 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2401.13178
openalex publication_date 2024/01/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Evaluating Large Language Models (LLMs) as general-purpose agents is essential for understanding their capabilities and facilitating their integration into practical applications. However, the evaluation process presents substantial challenges. A primary obstacle is the benchmarking of agent performance across diverse scenarios within a unified framework, especially in maintaining partially-observable environments and ensuring multi-round interactions. Moreover, current evaluation frameworks mostly focus on the final success rate, revealing few insights during the process and failing to provide a deep understanding of the model abilities. To address these challenges, we introduce AgentBoard, a pioneering comprehensive benchmark and accompanied open-source evaluation framework tailored to analytical evaluation of LLM agents. AgentBoard offers a fine-grained progress rate metric that captures incremental advancements as well as a comprehensive evaluation toolkit that features easy assessment of agents for multi-faceted analysis. This not only sheds light on the capabilities and limitations of LLM agents but also propels the interpretability of their performance to the forefront. Ultimately, AgentBoard serves as a step towards demystifying agent behaviors and accelerating the development of stronger LLM agents.
Cited by
- Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation
- SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows
- Fed-SE: Federated Self-Evolution for Privacy-Constrained Multi-Environment LLM Agents
- RefineBench: Evaluating Refinement Capability of Language Models via Checklists
- LeechHijack: Covert Computational Resource Exploitation in Intelligent Agent Systems
- STRIDE: A Systematic Framework for Selecting AI Modalities -- Agentic AI, AI Assistants, or LLM Calls
- SkeletonAgent: An Agentic Interaction Framework for Skeleton-based Action Recognition
- Conversational No-code, Multi-agentic Disease Module Identification and Drug Repurposing Prediction with ChatDRex
- AutoTool: Efficient Tool Selection for Large Language Model Agents
- TPS-Bench: Evaluating AI Agents' Tool Planning & Scheduling Abilities in Compounding Tasks
- Structured Uncertainty guided Clarification for LLM Agents
- Interaction as Intelligence Part II: Asynchronous Human-Agent Rollout for Long-Horizon Task Training
- The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
- WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing
- OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation
- Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
- GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
- AgentCaster: Reasoning-Guided Tornado Forecasting
- FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasks
- A Large-Language-Model Assisted Automated Scale Bar Detection and Extraction Framework for Scanning Electron Microscopic Images
- FlowSearch: Advancing deep research with dynamic structured knowledge flow
- Drift No More? Context Equilibria in Multi-Turn LLM Interactions
- A2Search: Ambiguity-Aware Question Answering with Reinforcement Learning
- AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework
- Agentic Services Computing
- PSG-Agent: Personality-Aware Safety Guardrail for LLM-based Agents
- Evaluating Agentic Bioinformatics through Function, Evidence, and Validation
- Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems
- Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs
- Offline Preference-Based Trajectory Evaluation
- Agentic AutoSurvey: Let LLMs Survey LLMs
- An Evaluation-Centric Paradigm for Scientific Visualization Agents
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- H2R: Hierarchical Hindsight Reflection for Multi-Task LLM Agents
- RAFFLES: Reasoning-based Attribution of Faults for LLM Systems
- FDABench: A Benchmark for Data Agents on Analytical Queries over Heterogeneous Data
- IndusGCC: A Data Benchmark and Evaluation Framework for GUI-Based General Computer Control in Industrial Automation
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- CardAIc-Agents: A Multimodal Framework with Hierarchical Adaptation for Cardiac Care Support
- What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles
- Intrinsic Memory Agents: Heterogeneous Multi-Agent LLM Systems through Structured Contextual Memory
- Galaxy: A Cognition-Centered Framework for Proactive, Privacy-Preserving, and Self-Evolving LLM Agents
- CoEx -- Co-evolving World-model and Exploration
- Evaluation and Benchmarking of LLM Agents: A Survey
- MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
- PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost
- FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain
- MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models
- A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
- NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability
- MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft
- EvoAgentX: An Automated Framework for Evolving Agentic Workflows
- Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation
- G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems
- KunLunBaizeRAG: Reinforcement Learning Driven Inference Performance Leap for Large Language Models
- Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories
- OmniReflect: Discovering Transferable Constitutions for LLM agents via Neuro-Symbolic Reflections
- Towards Pervasive Distributed Agentic Generative AI -- A State of The Art
- AgentSwift: Efficient LLM Agent Design via Value-guided Hierarchical Search
- From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems
- LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback
- MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
- A Practical Approach for Building Production-Grade Conversational Agents with Workflow Graphs
- Make Planning Research Rigorous Again!
- ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
- AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems
- GABench: A Comprehensive Benchmark for Evaluating LLM Agents on Graph Analysis Tasks
- Runaway is Ashamed, But Helpful: On the Early-Exit Behavior of Large Language Model-based Agents in Embodied Environments
- USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents
- ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
- MAPS: A Multilingual Benchmark for Agent Performance and Security
- ToolSpectrum : Towards Personalized Tool Utilization for Large Language Models
- Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges
- Learning to Play Like Humans: A Framework for LLM Adaptation in Interactive Fiction Games
- Retrospex: Language Agent Meets Offline Reinforcement Learning Critic
YC-Bench: Benchmarking AI Agents for Long-Term Planning and Consistent Execution- Exploration and Exploitation Errors Are Measurable for Language Model Agents
- Graph-based Agent Memory: Taxonomy, Techniques, and Applications
- UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
- Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use
- Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking
- SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model
- What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema
- Skill Discovery for Software Scripting Automation via Offline Simulations with LLMs
- Spec Kit Agents: Context-Grounded Agentic Workflows
- Auto-SLURP: A Benchmark Dataset for Evaluating Multi-Agent Frameworks in Smart Personal Assistant
- Verifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents
- PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
- Exploring Expert Failures Improves LLM Agent Tuning
- Matching Matters: A Fair Quality-Efficiency Benchmark for Command-Line Agents
- Where Agent Frameworks Fall Short: Examining Functional Challenges and Usability Concerns
- Breaking the Data Barrier -- Building GUI Agents Through Task Generalization
Related