AgentBench: Evaluating LLMs as Agents
2023/08/07 by Xiao Liu, Liu, Xiao, Hao Yu +41 · 133 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2308.03688
Abstract
The potential of Large Language Model (LLM) as agents has been widely acknowledged recently. Thus, there is an urgent need to quantitatively evaluate LLMs as agents on challenging tasks in interactive environments. We present AgentBench, a multi-dimensional benchmark that consists of 8 distinct environments to assess LLM-as-Agent's reasoning and decision-making abilities. Our extensive test over \num API-based and open-sourced (OSS) LLMs shows that, while top commercial LLMs present a strong ability of acting as agents in complex environments, there is a significant disparity in performance between them and many OSS competitors that are no larger than 70B. We identify the typical reasons of failures in environments and LLMs, showing that poor long-term reasoning, decision-making, and instruction following abilities are the main obstacles for developing usable LLM agents. Improving instruction following and training on high quality multi-round alignment data could improve agent performance. And different from existing assumptions, training on code present ambivalent impacts on different agent tasks. Datasets, environments, and an integrated evaluation package for AgentBench are released at https://github.com/THUDM/AgentBench.
Cited by
- Benchmark for Planning and Control with Large Language Model Agents: Blocksworld with Model Context Protocol
- Nested Browser-Use Learning for Agentic Information Seeking
- ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
- HiSciBench: A Hierarchical Multi-disciplinary Benchmark for Scientific Intelligence from Reading to Discovery
- E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios
- SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows
- SymStep: Symbolic Step Verification for Logical Reasoning
- HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- When Thinking Before Retrieval Hurts: TraceBound Diagnostics for Adaptive Knowledge-Graph Retrieval
- PeopleSearchBench: A Multi-Dimensional Benchmark for Evaluating AI-Powered People Search Platforms
- Multi-Agent LLM Committees for Autonomous Software Beta Testing
- Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
- Meta-RL Induces Exploration in Language Agents
- AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
- Small Language Models for Efficient Agentic Tool Calling: Outperforming Large Models with Targeted Fine-tuning
- City Navigation in the Wild: Exploring Emergent Navigation from Web-Scale Knowledge in MLLMs
- CangLing-KnowFlow: A Unified Knowledge-and-Flow-fused Agent for Comprehensive Remote Sensing Applications
- Towards a Science of Scaling Agent Systems
- AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration
- Can AI autonomously build, operate, and use the entire data stack?
- The Evolution of Agentic AI in Cybersecurity: From Single LLM Reasoners to Multi-Agent Systems and Autonomous Pipelines
- LoopBench: Discovering Emergent Symmetry Breaking Strategies with LLM Swarms
- TDAG: A multi-agent framework based on dynamic Task Decomposition and Agent Generation
- Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction
- SEAL: Self-Evolving Agentic Learning for Conversational Question Answering over Knowledge Graphs
- PPTArena: A Benchmark for PowerPoint Editing
- CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
- TinyLLM: Evaluation and Optimization of Small Language Models for Agentic Tasks on Edge Devices
- Benchmarking In-context Experiential Learning Through Repeated Product Recommendations
- OpenApps: Simulating Environment Variations to Measure UI-Agent Reliability
- Cross-LLM Generalization of Behavioral Backdoor Detection in AI Agent Supply Chains
- CostNav: A Navigation Benchmark for Real-World Economic-Cost Evaluation of Physical AI Agents
- ASTRA: Agentic Steerability and Risk Assessment Framework
- Why Do Language Model Agents Whistleblow?
- Fast LLM Post-training via Decoupled and Fastest-of-N Speculation
- Agent-R1: Training Powerful LLM Agents with End-to-End Reinforcement Learning
- TPS-Bench: Evaluating AI Agents' Tool Planning & Scheduling Abilities in Compounding Tasks
- UpBench: A Dynamically Evolving Real-World Labor-Market Agentic Benchmark Framework Built for Human-Centric AI
- ToolMind Technical Report: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset
- How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
- Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems
- VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- What's the next frontier for Data-centric AI? Data Savvy Agents
- AReaL-Hex: Accommodating Asynchronous RL Training over Heterogeneous GPUs
- Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
- Delegated Authorization for Agents Constrained to Semantic Task-to-Scope Matching
- OracleAgent: A Multimodal Reasoning Agent for Oracle Bone Script Research
- The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
- Agentic AI: A Comprehensive Survey of Architectures, Applications, and Future Directions
- RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents
- Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft
- Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
- APTBench: Benchmarking Agentic Potential of Base LLMs During Pre-Training
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- Unifying Large Language Models and Knowledge Graphs: A Roadmap
- DaMo: Data Mixing Optimizer in Fine-tuning Multimodal LLMs for Mobile Phone Agents
- WebGraphEval: Multi-Turn Trajectory Evaluation for Web Agents using Graph Representation
- AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- EvoSyn: Generalizable Evolutionary Data Synthesis for Verifiable Learning
- ProtocolBench: Which LLM MultiAgent Protocol to Choose?
- Static Sandboxes Are Inadequate: Modeling Societal Complexity Requires Open-Ended Co-Evolution in LLM-Based Multi-Agent Simulations
- Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
- GOAT: A Training Framework for Goal-Oriented Agent with Tools
- Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
- Failure-Driven Workflow Refinement
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces
- Can RL Improve Generalization of LLM Agents? An Empirical Study
- What Is Your Agent's GPA? A Framework for Evaluating Agent Goal-Plan-Action Alignment
- Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation
- Code Agent can be an End-to-end System Hacker: Benchmarking Real-world Threats of Computer-use Agent
- MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline
- BIRD-INTERACT: Re-imagining Text-to-SQL Evaluation for Large Language Models via Lens of Dynamic Interactions
- Where Did It All Go Wrong? A Hierarchical Look into Multi-Agent Error Attribution
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- MacroBench: A Novel Testbed for Web Automation Scripts via Large Language Models
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- Demystifying deep search: a holistic evaluation with hint-free multi-hop questions and factorised metrics
- The Social Laboratory: A Psychometric Framework for Multi-Agent LLM Evaluation
- Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs
- MEMTRACK: Evaluating Long-Term Memory and State Tracking in Multi-Platform Dynamic Agent Environments
- A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
- A Measurement Study of Model Context Protocol Ecosystem
- Agentic Services Computing
- The Matthew Effect of AI Programming Assistants: A Hidden Bias in Software Evolution
- BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics
- Can AI Perceive Physical Danger and Intervene?
- Regulating the Agency of LLM-based Agents
- Automotive-ENV: Benchmarking Multimodal Agents in Vehicle Interface Systems
- CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineering
- What Do LLM Agents Do When Left Alone? Evidence of Spontaneous Meta-Cognitive Patterns
- Exploration with Foundation Models: Capabilities, Limitations, and Hybrid Approaches
- Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness
- DeepResearch Agent System
- Scaling LLM-Driven Multi-Agent Systems: Design Principles and Architectural Scalability Analysis
- ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents
- Evaluating LLM Agents on Automated Software Analysis Tasks
- MSCoRe: A Benchmark for Multi-Stage Collaborative Reasoning in LLM Agents
- An Evaluation-Centric Paradigm for Scientific Visualization Agents
- Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
- GEM-Bench: A Benchmark for Ad-Injected Response Generation within Generative Engine Marketing
- TGPO: Tree-Guided Preference Optimization for Robust Web Agent Reinforcement Learning
- Who is Introducing the Failure? Automatically Attributing Failures of Multi-Agent Systems via Spectrum Analysis
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- PosterGen: Aesthetic-Aware Paper-to-Poster Generation via Multi-Agent LLMs
- Redefining Website Fingerprinting Attacks With Multiagent LLMs
- AI Wellbeing
- Agents of Discovery
- Internet 3.0: Architecture for a Web-of-Agents with it's Algorithm for Ranking Agents
- KubeIntellect: A Modular LLM-Orchestrated Agent Framework for End-to-End Kubernetes Management
- Can Large Language Models Master Complex Card Games?
- Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making
- A Study on the Framework for Evaluating the Ethics and Trustworthiness of Generative AI
- Transforming Agency. On the mode of existence of Large Language Models
- NiceWebRL: a Python library for human subject experiments with reinforcement learning environments
- CataractSurg-80K: Knowledge-Driven Benchmarking for Structured Reasoning in Ophthalmic Surgery Planning
- Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
- SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling
- AI Agentic Programming: A Survey of Techniques, Challenges, and Opportunities
- Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs
- GVGAI-LLM: Evaluating Large Language Model Agents with Infinite Games
- Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
- OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
- Pro2Guard: Proactive Runtime Enforcement of LLM Agent Safety via Probabilistic Model Checking
- PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning
Related