GAIA: a benchmark for General AI Assistants
2023/11/21 by Grégoire Mialon, Mialon, Grégoire, Clémentine Fourrier +9 · 6 voices · 406 citations
Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Artificial intelligence #Benchmark (surveying) #Computer science #Explainable Artificial Intelligence (XAI) #Milestone #Plug-in #Programming language #Robustness (evolution) #Set (abstract data type) #Topic Modeling #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2311.12983
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/11/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We introduce GAIA, a benchmark for General AI Assistants that, if solved, would represent a milestone in AI research. GAIA proposes real-world questions that require a set of fundamental abilities such as reasoning, multi-modality handling, web browsing, and generally tool-use proficiency. GAIA questions are conceptually simple for humans yet challenging for most advanced AIs: we show that human respondents obtain 92% vs. 15% for GPT-4 equipped with plugins. This notable performance disparity contrasts with the recent trend of LLMs outperforming humans on tasks requiring professional skills in e.g. law or chemistry. GAIA's philosophy departs from the current trend in AI benchmarks suggesting to target tasks that are ever more difficult for humans. We posit that the advent of Artificial General Intelligence (AGI) hinges on a system's capability to exhibit similar robustness as the average human does on such questions. Using GAIA's methodology, we devise 466 questions and their answer. We release our questions while retaining answers to 300 of them to power a leader-board available at https://huggingface.co/gaia-benchmark.
Cited by
- InteractComp: Evaluating Search Agents With Ambiguous Queries
- Agentic Evaluation of Copyright Law Compliance
- Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI
- DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents
- From Mind to Machine: The Rise of Manus AI as a Fully Autonomous Digital Agent
- FedAgentKE: Federated Semantic Knowledge Evolution for Heterogeneous Agents
- ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability
- The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation
- In-the-Flow Agentic System Optimization for Effective Planning and Tool Use
- CEO-Bench: Can Agents Play the Long Game?
- ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management
- AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
- Efficient Multi-round LLM Inference over Disaggregated Serving
- AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
- Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness
- WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
- MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking
- Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains
- Where Does Agent Reliability Come From? A Cross-Benchmark Decomposition of Verification Loops, Specialist Models, and Scaffolding in a Production Enterprise Agent
- OTAP: Structure-Aware Optimal Transport for Evaluating Planning and Execution in Agent Trajectories
- RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications
- DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments
- Do Generative AI Assistants Respect robots.txt? Tracing Web Access Beyond Visible Answers
- PCTD: Preference-Guided Counterfactual Task Decomposition for Agent Tool Retrieval
- Understanding Agent-Reactive Bugs at the Model-Harness Boundary: An Empirical Study of LLM Agent Issue Reports
- Self-Aware Recursively Self-Improving Agents for Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture
- SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
- Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain
- Autonomous Topology Mutation: Safe Runtime Restructuring for Multi-Agent LLM Systems with Capability, State, and Shadow Invariants
- Meta-Reinforcement Learning with Self-Reflection for Agentic Search
- Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem
- Remote Labor Index: Measuring AI Automation of Remote Work
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
- Magentic-UI: Towards Human-in-the-loop Agentic Systems
- Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce
- Agentic Entropy-Balanced Policy Optimization
- Nested Browser-Use Learning for Agentic Information Seeking
- Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents
- Video-BrowseComp: Benchmarking Agentic Video Research on Open Web
- FoldAct: Efficient and Stable Context Folding for Long-Horizon Search Agents
- Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation
- Delegation Intelligence in Deep Search: A Controllable Framework for Disentangled Capability Diagnosis
- E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios
- ScholarSearch: Benchmarking Scholar Searching Ability of LLMs
- SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows
- HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems
- LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
- LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
- Step-DeepResearch Technical Report
- MemEvolve: Meta-Evolution of Agent Memory Systems
- OpenView: Empowering MLLMs with Out-of-view VQA
- DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
- SCOPE: Prompt Evolution for Enhancing Agent Effectiveness
- ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research
- AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression
- Benchmarking the Generality of Vision-Language-Action Models
- Unifying Dynamic Tool Creation and Cross-Task Experience Sharing through Cognitive Memory Architecture
- FutureWeaver: Planning Test-Time Compute for Multi-Agent Systems with Modularized Collaboration
- The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
- EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce
- DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
- How Well Does Agent Development Reflect Real-World Work?
- CARL: Criticality-Aware Agentic Reinforcement Learning
- PPTArena: A Benchmark for PowerPoint Editing
- Think in Parallel, Answer as One: Logit Averaging for Open-Ended Reasoning
- LeechHijack: Covert Computational Resource Exploitation in Intelligent Agent Systems
- How Far Are We from Genuinely Useful Deep Research Agents?
- AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets
- CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
- ShoppingComp: Are LLMs Really Ready for Your Shopping Cart?
- OpenApps: Simulating Environment Variations to Measure UI-Agent Reliability
- CostNav: A Navigation Benchmark for Real-World Economic-Cost Evaluation of Physical AI Agents
- Fara-7B: An Efficient Agentic Model for Computer Use
- PRInTS: Reward Modeling for Long-Horizon Information Seeking
- AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning
- RhinoInsight: Improving Deep Research through Control Mechanisms for Model Behavior and Context
- ASTRA: Agentic Steerability and Risk Assessment Framework
- M3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
- TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots
- Computer-Use Agents as Judges for Generative User Interface
- Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
- MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
- Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning
- Environment Scaling for Interactive Agentic Experience Collection: A Survey
- DigiData: Training and Evaluating General-Purpose Mobile Control Agents
- IterResearch: Rethinking Long-Horizon Agents via Markovian State Reconstruction
- The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents
- A Hierarchical Imprecise Probability Approach to Reliability Assessment of Large Language Models
- Delegated Authorization for Agents Constrained to Semantic Task-to-Scope Matching
- Decomposition-Enhanced Training for Post-Hoc Attributions In Language Models
- Completion ≠ Collaboration: Scaling Collaborative Effort with Agents
- Agentic AI: A Comprehensive Survey of Architectures, Applications, and Future Directions
- Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing
- OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation
- Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
- Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI
- Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios
- TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- Agents' Last Exam
- AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
- DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent Behaviors
- Automated Composition of Agents: A Knapsack Approach for Agentic Component Selection
- MAS-ProVe: Understanding the Process Verification of Multi-Agent Systems
- Model-Document Protocol for AI Search
- Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
- Tongyi DeepResearch Technical Report
- AgentFold: Long-Horizon Web Agents with Proactive Context Management
- ParallelMuse: Agentic Parallel Thinking for Deep Information Seeking
- WebLeaper: Empowering Efficiency and Efficacy in WebAgent via Enabling Info-Rich Seeking
- AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis
- Repurposing Synthetic Data for Fine-grained Search Agent Supervision
- What Limits Agentic Systems Efficiency?
- APTBench: Benchmarking Agentic Potential of Base LLMs During Pre-Training
- MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP Tools
- Alita-G: Self-Evolving Generative Agent for Agent Generation
- DeepAgent: A General Reasoning Agent with Scalable Toolsets
- Co-Sight: Enhancing LLM-Based Agents via Conflict-Aware Meta-Verification and Trustworthy Reasoning with Structured Facts
- Rethinking the Design of Reinforcement Learning-Based Deep Research Agents
- DeepWideSearch: Benchmarking Depth and Width in Agentic Information Seeking
- ETOM: A Five-Level Benchmark for Evaluating Tool Orchestration within the MCP Ecosystem
- DaMo: Data Mixing Optimizer in Fine-tuning Multimodal LLMs for Mobile Phone Agents
- TheMCPCompany: Creating General-purpose Agents with Task-specific Tools
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- Offline Policy Evaluation of Multi-Turn LLM Health Coaching with Real Users
- ProtocolBench: Which LLM MultiAgent Protocol to Choose?
- UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- Explore to Evolve: Scaling Evolved Aggregation Logic via Proactive Online Exploration for Deep Research Agents
- FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis
- Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms
- EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
- ConsintBench: Evaluating Language Models on Real-World Consumer Intent Understanding
- Laminar: A Scalable Asynchronous RL Post-Training Framework
- ResearStudio: A Human-Intervenable Framework for Building Controllable Deep-Research Agents
- Deep Research Brings Deeper Harm
- ACADREASON: Exploring the Limits of Reasoning Models with Academic Research Problems
- A2FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning
- PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
- StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
- BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
- How can we assess human-agent interactions? Case studies in software agent design
- Agentic-KGR: Co-evolutionary Knowledge Graph Construction through Multi-Agent Reinforcement Learning
- COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context
- MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
- FlowSearch: Advancing deep research with dynamic structured knowledge flow
- Beyond Turn Limits: Training Deep Search Agents with Dynamic Context Window
- DISCO: Diversifying Sample Condensation for Efficient Model Evaluation
- Understanding DeepResearch via Reports
- What Is Your Agent's GPA? A Framework for Evaluating Agent Goal-Plan-Action Alignment
- WebDART: Dynamic Decomposition and Re-planning for Complex Web Tasks
- Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
- MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline
- Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification
- LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation
- Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports
- BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
- TaskCraft: Automated Generation of Agentic Tasks
- Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- JoyAgent-JDGenie: Technical Report on the GAIA
- Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
- From Factoid Questions to Data Product Requests: Benchmarking Data Product Discovery over Tables and Text
- Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs
- When Hallucination Costs Millions: Benchmarking AI Agents in High-Stakes Adversarial Financial Markets
- SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA
- Where LLM Agents Fail and How They can Learn From Failures
- InfoAgent: Advancing Autonomous Information-Seeking Agents
- Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution
- Scaling Generalist Data-Analytic Agents
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
- Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents
- Demystifying the Lifecycle of Failures in Platform-Orchestrated Agentic Workflows
- SafeSearch: Automated Red-Teaming for the Safety of LLM-Based Search Agents
- CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems
- Tree Search for LLM Agent Reinforcement Learning
- CORE: Full-Path Evaluation of LLM Agents Beyond Final State
- Agent Harness Distillation: Inference-Time Harness Extraction and Exploitation in Autonomous Multi-Agent Systems
- ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents
- A Policy-Driven Runtime Layer for Agentic LLM Serving
- Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
- Designing for Doubt: The Case for Informed Abstention in Autonomous Agents
- The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
- A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
- MSCoRe: A Benchmark for Multi-Stage Collaborative Reasoning in LLM Agents
- ARE: Scaling Up Agent Environments and Evaluations
- An Evaluation-Centric Paradigm for Scientific Visualization Agents
- AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production
- Who is Introducing the Failure? Automatically Attributing Failures of Multi-Agent Systems via Spectrum Analysis
- Aegis: Automated Error Generation and Attribution for Multi-Agent Systems
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning
- WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research
- WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
- ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization
- Scaling Agents via Continual Pre-training
- Anemoi: A Semi-Centralized Multi-agent System Based on Agent-to-Agent Communication MCP server from Coral Protocol
- Difficulty-Aware Agentic Orchestration for Query-Specific Multi-Agent Workflows
- Abduct, Act, Predict: Scaffolding Causal Inference for Automated Failure Attribution in Multi-Agent Systems
- RAFFLES: Reasoning-based Attribution of Faults for LLM Systems
- Reinforcement Learning Foundations for Deep Research Systems: A Survey
- WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
- SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
- Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects
- AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
- FDABench: A Benchmark for Data Agents on Analytical Queries over Heterogeneous Data
- Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- Physics Supernova: AI Agent Matches Elite Gold Medalists at IPhO 2025
- HiVA: Self-organized Hierarchical Variable Agent via Goal-driven Semantic-Topological Evolution
- How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on τ-bench
- AWorld: Orchestrating the Training Recipe for Agentic AI
- DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
- Foundational Design Principles and Patterns for Building Robust and Adaptive GenAI-Native Systems
- MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
- ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction
- Deep Research: A Survey of Autonomous Research Agents
- Fast, Slow, and Tool-augmented Thinking for LLMs: A Review
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
- MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
- ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
- BrowseMaster: Towards Scalable Web Browsing via Tool-Augmented Programmatic Agent Pair
- WideSearch: Benchmarking Agentic Broad Info-Seeking
- Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL
- SKATE, a Scalable Tournament Eval: Weaker LLMs differentiate between stronger ones using verifiable challenges
- WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging
- ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents
- Characterizing Deep Research: A Benchmark and Formal Definition
- VeriGUI: Verifiable Long-Chain GUI Dataset
- HealthFlow: A Self-Evolving AI Agent with Meta Planning for Autonomous Healthcare Research
- A Survey of LLM-based Deep Search Agents: Paradigm, Optimization, Evaluation, and Challenges
- WebDS: An End-to-End Benchmark for Web-based Data Science
- MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning
- Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
- Agentic Reinforced Policy Optimization
- OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
- WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization
- Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments
- Agent Identity Evals: Measuring Agentic Identity
- FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain
- Beyond Isolated Dots: Benchmarking Structured Table Construction as Deep Knowledge Extraction
- Aime: Towards Fully-Autonomous Multi-Agent Framework
- Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs
- Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
- Optimizing Sequential Multi-Step Tasks with Parallel LLM Agents
- AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs
- Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
- Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights
- Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving
- Deep Research Comparator: A Platform For Fine-grained Human Annotations of Deep Research Agents
- Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework
- ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL
- Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
- Autonomous Repair for Multi-Agent Systems via Monte-Carlo Tree Search
- Hijacking JARVIS: Benchmarking Mobile GUI Agents against Unprivileged Third Parties
- Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications
- CodeAgents: A Token-Efficient Framework for Codified Multi-Agent Reasoning in LLMs
- APPO: Agentic Procedural Policy Optimization
- EvoAgentX: An Automated Framework for Evolving Agentic Workflows
- Establishing Best Practices for Building Rigorous Agentic Benchmarks
- HiRA: A Hierarchical Reasoning Framework for Decoupled Planning and Execution in Deep Search
- WebSailor: Navigating Super-human Reasoning for Web Agent
- The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?
- La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America
- DABstep: Data Agent Benchmark for Multi-step Reasoning
- The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
- Compound AI Systems Optimization: A Survey of Methods, Challenges, and Future Directions
- Deep Research Agents: A Systematic Examination And Roadmap
- Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents
- SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
- OAgents: An Empirical Study of Building Effective Agents
- A Vision for Geo-Temporal Deep Research Systems: Towards Comprehensive, Transparent, and Reproducible Geo-Temporal Information Synthesis
- Code Reasoning for Software Engineering Tasks: A Survey and A Call to Action
- Scaling Test-time Compute for LLM Agents
- PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
- Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
- OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
- FinDeepIndicator: Benchmarking Deep Research Agents in End-to-End Financial Indicator Construction
- macOSWorld: A Multilingual Interactive Benchmark for GUI Agents
- Coding Agents with Multimodal Browsing are Generalist Problem Solvers
- EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving
- AI-Native Brand Identity: From Visual Recognition to Cryptographic Verification
- WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
- Divide, Optimize, Merge: Fine-Grained LLM Agent Optimization at Scale
- Deep Research Bench: Evaluating AI Web Research Agents
- ThinkGeo: Evaluating Tool-Augmented Agents for Remote Sensing Tasks
- OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
- WebDancer: Towards Autonomous Information Seeking Agency
- WebCoT: Enhancing Web Agent Reasoning by Reconstructing Chain-of-Thought in Reflection, Branching, and Rollback
- Training LLM-Based Agents with Synthetic Self-Reflected Trajectories and Partial Masking
- Multi-View Encoders for Performance Prediction in LLM-Based Agentic Workflows
- Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution
- On Path to Multimodal Historical Reasoning: HistBench and HistAgent
- DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research
- HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses
- SearchMaster: Grounded and Regulated Self-Play for Search Agents
- FedWorld: Scope-Aware Federation of Agent World Models
- The Real Barrier to LLM Agent Usability is Agentic ROI
- From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution
- T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning
- MCP-RADAR: A Multi-Dimensional Benchmark for Evaluating Tool Use Capabilities in Large Language Models
- Manalyzer: End-to-end Automated Meta-analysis with Multi-agent System
- Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
- O2-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering
- SimpleDeepSearcher: Deep Information Seeking via Web-Powered Reasoning Trajectory Synthesis
- Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce
- LLM-Powered AI Agent Systems and Their Applications in Industry
- MASLab: A Unified and Comprehensive Codebase for LLM-based Multi-Agent Systems
- lmgame-Bench: How Good are LLMs at Playing Games?
- MAPS: A Multilingual Benchmark for Agent Performance and Security
- GISAgentBench: A Practitioner-Sourced Benchmark for Evaluating LLM Agents on GIS Tasks
- ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions
- MedBrowseComp: Benchmarking Medical Deep Research and Computer Use
- TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents
- Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges
- G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution
- AgentSLABench: Evaluating and Benchmarking Agentic Systems Under Resource Constraints
- Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation
- Adversarial Attacks in Multi-Agent LLM Pipelines: Unveiling Structural Vulnerabilities in Agentic AI Architectures
- DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment
- AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
- TRAIL: Trace Reasoning and Agentic Issue Localization
- FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
- Internet of Agents: Fundamentals, Applications, and Challenges
- EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data
- HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
- SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization
YC-Bench: Benchmarking AI Agents for Long-Term Planning and Consistent Execution- R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning
- Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
- EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
- From Model Scaling to System Scaling: Scaling the Harness in Agentic AI
- The Bystander Effect in Multi-Agent Reasoning: Quantifying Cognitive Loafing in Collaborative Interactions
- Personalized Deep Research: A User-Centric Framework, Dataset, and Hybrid Evaluation for Knowledge Discovery
- Graph-based Agent Memory: Taxonomy, Techniques, and Applications
- UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
- PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents
- Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use
- Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking
- Anticipatory Planning for Multimodal AI Agents
- IPO Finance Agent: Benchmark of LLM Financial Analysts Beyond Finance Agent v2, with Automated Rubric Generation, on the SpaceX (SPCX) IPO
- OpenBioRQ: Unsolved Biomedical Research Questions for Agents
- MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
- GRPO Does Not Close the Multi-Agent Coordination Gap
- $OneMillion-Bench: How Far are Language Agents from Human Experts?
- From Failed Trajectories to Reliable LLM Agents: Diagnosing and Repairing Harness Flaws
- Architecting Trust in Artificial Epistemic Agents
- Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
- VeRO: A Harness for Agents to Optimize Agents
- Benchmark Test-Time Scaling of General LLM Agents
- AgentConductor: Topology Evolution for Multi-Agent Competition-Level Code Generation
- Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
- SWE Context Bench: A Benchmark for Context Learning in Coding
- Paper2Rebuttal: A Multi-Agent Framework for Transparent Author Response Assistance
- WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks
- Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
- What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema
- WebThinker: Empowering Large Reasoning Models with Deep Research Capability
- PIVOT: Bridging Planning and Execution in LLM Agents via Trajectory Refinement
- Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?
- OpenThoughts-Agent: Data Recipes for Agentic Models
- Skill Discovery for Software Scripting Automation via Offline Simulations with LLMs
- From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
- Toward a Science of Intent: Closure Gaps and Delegation Envelopes for Open-World AI Agents
- Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent
- Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
- Beyond Parallel Sampling: Diverse Query Initialization for Agentic Search
- Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
- MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models
- GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration
- BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese
- Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
- GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows
- Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning
- InfoSeeker: A Scalable Hierarchical Parallel Agent Framework for Web Information Seeking
- Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests
- Beyond Rule-Based Workflows: An Information-Flow-Orchestrated Multi-Agents Paradigm via Agent-to-Agent Communication from CORAL
- DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation
- MemoBrain: Executive Memory as an Agentic Brain for Reasoning
- DeepSynth-Eval: Objectively Evaluating Information Consolidation in Deep Survey Writing
- WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
- ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
- Auditing the Ethical Logic of Generative AI Models
- JITServe: SLO-aware LLM Serving with Imprecise Request Information
- Architectural Implications of Agentic AI Workflows
- FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation
- MatrAIx: Simulating the World with 8.3 Billion Persona Agents
- FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables
- WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model
- Synergizing RAG and Reasoning: A Systematic Review
- LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
- MARFT: Multi-Agent Reinforcement Fine-Tuning
- Planet as a Brain: Towards Internet of AgentSites based on AIOS Server
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- A Two-Tier Perspective on Inference-Time Parallelism in Multi-Agent LLM Systems
- EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents
- SP-Mind: An Autonomous Reasoning Agent for Spatial Proteomics Analysis
- COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
Discussions
- Meta: Gaia - A Benchmark for General AI Assistants [hn, 36 points, 8 comments]
- haven't seen any of the AI opinion havers on here discussing GAIA yet arxiv.org/abs/2311.12983 [bsky, 1 points, 0 comments]
- GAIA challenges AI with tasks easy for humans but tough for AI, showing a 92% success rate for humans versus 15% for GPT-4. The benchmark includes 466 questions designed to test fundamental abilities [bsky, 0 points, 0 comments]
- tl;dr — AI is dumb. arxiv.org/abs/2311.12983 [bsky, 0 points, 0 comments]
- Pre-print of research paper arxiv.org/abs/2311.12983 [bsky, 0 points, 0 comments]
- Benchmark for General AI Assistants [Mialon+, 2023] GAIA evaluates AI's fundamental abilities, such as web search, image understanding, and file reading, rather than expert knowledge. While humans ach [bsky, 0 points, 0 comments]
Related