GAIA: a benchmark for General AI Assistants
2023/11/21 by Grégoire Mialon, Mialon, Grégoire, Clémentine Fourrier +9 · 6 voices · 238 citations
Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI) #Topic Modeling #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2311.12983
openalex publication_date 2023/11/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We introduce GAIA, a benchmark for General AI Assistants that, if solved, would represent a milestone in AI research. GAIA proposes real-world questions that require a set of fundamental abilities such as reasoning, multi-modality handling, web browsing, and generally tool-use proficiency. GAIA questions are conceptually simple for humans yet challenging for most advanced AIs: we show that human respondents obtain 92% vs. 15% for GPT-4 equipped with plugins. This notable performance disparity contrasts with the recent trend of LLMs outperforming humans on tasks requiring professional skills in e.g. law or chemistry. GAIA's philosophy departs from the current trend in AI benchmarks suggesting to target tasks that are ever more difficult for humans. We posit that the advent of Artificial General Intelligence (AGI) hinges on a system's capability to exhibit similar robustness as the average human does on such questions. Using GAIA's methodology, we devise 466 questions and their answer. We release our questions while retaining answers to 300 of them to power a leader-board available at https://huggingface.co/gaia-benchmark.
Cited by
- InteractComp: Evaluating Search Agents With Ambiguous Queries
- Agentic Evaluation of Copyright Law Compliance
- Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI
- DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents
- From Mind to Machine: The Rise of Manus AI as a Fully Autonomous Digital Agent
- FedAgentKE: Federated Semantic Knowledge Evolution for Heterogeneous Agents
- ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability
- The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation
- In-the-Flow Agentic System Optimization for Effective Planning and Tool Use
- CEO-Bench: Can Agents Play the Long Game?
- ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management
- AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
- Efficient Multi-round LLM Inference over Disaggregated Serving
- AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
- Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness
- WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
- MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking
- Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains
- Where Does Agent Reliability Come From? A Cross-Benchmark Decomposition of Verification Loops, Specialist Models, and Scaffolding in a Production Enterprise Agent
- Otap:Structure-Aware Optimal Transport for Evaluating Planning and Execution in Agent Trajectories
- RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications
- DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments
- Do Generative AI Assistants Respect robots.txt? Tracing Web Access Beyond Visible Answers
- PCTD: Preference-Guided Counterfactual Task Decomposition for Agent Tool Retrieval
- Understanding Agent-Reactive Bugs at the Model-Harness Boundary: An Empirical Study of LLM Agent Issue Reports
- Self-Aware Recursively Self-Improving Agents for Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture
- SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
- Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain
- Autonomous Topology Mutation: Safe Runtime Restructuring for Multi-Agent LLM Systems with Capability, State, and Shadow Invariants
- Meta-Reinforcement Learning with Self-Reflection for Agentic Search
- Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem
- Remote Labor Index: Measuring AI Automation of Remote Work
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
- Magentic-UI: Towards Human-in-the-loop Agentic Systems
- Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce
- Agentic Entropy-Balanced Policy Optimization
- Nested Browser-Use Learning for Agentic Information Seeking
- Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents
- Video-BrowseComp: Benchmarking Agentic Video Research on Open Web
- FoldAct: Efficient and Stable Context Folding for Long-Horizon Search Agents
- Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation
- Delegation Intelligence in Deep Search: A Controllable Framework for Disentangled Capability Diagnosis
- E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios
- SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows
- HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems
- LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
- LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
- Step-DeepResearch Technical Report
- MemEvolve: Meta-Evolution of Agent Memory Systems
- OpenView: Empowering MLLMs with Out-of-view VQA
- DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
- SCOPE: Prompt Evolution for Enhancing Agent Effectiveness
- ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research
- Multi-Agent Collaborative Framework for Intelligent IT Operations: An AOI System with Context-Aware Compression and Dynamic Task Scheduling
- Benchmarking the Generality of Vision-Language-Action Models
- Unifying Dynamic Tool Creation and Cross-Task Experience Sharing through Cognitive Memory Architecture
- FutureWeaver: Planning Test-Time Compute for Multi-Agent Systems with Modularized Collaboration
- The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
- EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce
- DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
- How Well Does Agent Development Reflect Real-World Work?
- CARL: Criticality-Aware Agentic Reinforcement Learning
- PPTArena: A Benchmark for PowerPoint Editing
- Think in Parallel, Answer as One: Logit Averaging for Open-Ended Reasoning
- LeechHijack: Covert Computational Resource Exploitation in Intelligent Agent Systems
- How Far Are We from Genuinely Useful Deep Research Agents?
- AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets
- CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
- ShoppingComp: Are LLMs Really Ready for Your Shopping Cart?
- OpenApps: Simulating Environment Variations to Measure UI-Agent Reliability
- CostNav: A Navigation Benchmark for Real-World Economic-Cost Evaluation of Physical AI Agents
- Fara-7B: An Efficient Agentic Model for Computer Use
- PRInTS: Reward Modeling for Long-Horizon Information Seeking
- AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning
- RhinoInsight: Improving Deep Research through Control Mechanisms for Model Behavior and Context
- ASTRA: Agentic Steerability and Risk Assessment Framework
- M3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
- TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots
- Computer-Use Agents as Judges for Generative User Interface
- Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
- MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
- Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning
- Environment Scaling for Interactive Agentic Experience Collection: A Survey
- DigiData: Training and Evaluating General-Purpose Mobile Control Agents
- IterResearch: Rethinking Long-Horizon Agents via Markovian State Reconstruction
- The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents
- A Hierarchical Imprecise Probability Approach to Reliability Assessment of Large Language Models
- Delegated Authorization for Agents Constrained to Semantic Task-to-Scope Matching
- Decomposition-Enhanced Training for Post-Hoc Attributions In Language Models
- Completion ≠ Collaboration: Scaling Collaborative Effort with Agents
- Agentic AI: A Comprehensive Survey of Architectures, Applications, and Future Directions
- Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing
- OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation
- Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
- Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI
- Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios
- TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- Agents' Last Exam
- AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
- DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent Behaviors
- Automated Composition of Agents: A Knapsack Approach for Agentic Component Selection
- MAS-ProVe: Understanding the Process Verification of Multi-Agent Systems
- Model-Document Protocol for AI Search
- Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
- Tongyi DeepResearch Technical Report
- AgentFold: Long-Horizon Web Agents with Proactive Context Management
- ParallelMuse: Agentic Parallel Thinking for Deep Information Seeking
- WebLeaper: Empowering Efficiency and Efficacy in WebAgent via Enabling Info-Rich Seeking
- AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis
- Repurposing Synthetic Data for Fine-grained Search Agent Supervision
- What Limits Agentic Systems Efficiency?
- APTBench: Benchmarking Agentic Potential of Base LLMs During Pre-Training
- MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP Tools
- Alita-G: Self-Evolving Generative Agent for Agent Generation
- DeepAgent: A General Reasoning Agent with Scalable Toolsets
- Co-Sight: Enhancing LLM-Based Agents via Conflict-Aware Meta-Verification and Trustworthy Reasoning with Structured Facts
- Rethinking the Design of Reinforcement Learning-Based Deep Research Agents
- DeepWideSearch: Benchmarking Depth and Width in Agentic Information Seeking
- MSC-Bench: A Rigorous Benchmark for Multi-Server Tool Orchestration
- DaMo: Data Mixing Optimizer in Fine-tuning Multimodal LLMs for Mobile Phone Agents
- TheMCPCompany: Creating General-purpose Agents with Task-specific Tools
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- Offline Policy Evaluation of Multi-Turn LLM Health Coaching with Real Users
- ProtocolBench: Which LLM MultiAgent Protocol to Choose?
- UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- Explore to Evolve: Scaling Evolved Aggregation Logic via Proactive Online Exploration for Deep Research Agents
- FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis
- Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms
- EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
- ConsintBench: Evaluating Language Models on Real-World Consumer Intent Understanding
- Laminar: A Scalable Asynchronous RL Post-Training Framework
- ResearStudio: A Human-Intervenable Framework for Building Controllable Deep-Research Agents
- Deep Research Brings Deeper Harm
- ACADREASON: Exploring the Limits of Reasoning Models with Academic Research Problems
- A2FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning
- PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
- StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
- BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
- How can we assess human-agent interactions? Case studies in software agent design
- Agentic-KGR: Co-evolutionary Knowledge Graph Construction through Multi-Agent Reinforcement Learning
- COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context
- MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
- FlowSearch: Advancing deep research with dynamic structured knowledge flow
- Beyond Turn Limits: Training Deep Search Agents with Dynamic Context Window
- DISCO: Diversifying Sample Condensation for Efficient Model Evaluation
- Understanding DeepResearch via Reports
- What Is Your Agent's GPA? A Framework for Evaluating Agent Goal-Plan-Action Alignment
- WebDART: Dynamic Decomposition and Re-planning for Complex Web Tasks
- Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
- MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline
- Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification
- LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation
- Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports
- BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
- Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- JoyAgent-JDGenie: Technical Report on the GAIA
- Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
- From Factoid Questions to Data Product Requests: Benchmarking Data Product Discovery over Tables and Text
- Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs
- When Hallucination Costs Millions: Benchmarking AI Agents in High-Stakes Adversarial Financial Markets
- SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA
- Where LLM Agents Fail and How They can Learn From Failures
- InfoAgent: Advancing Autonomous Information-Seeking Agents
- Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution
- Scaling Generalist Data-Analytic Agents
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
- Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents
- Demystifying the Lifecycle of Failures in Platform-Orchestrated Agentic Workflows
- SafeSearch: Automated Red-Teaming for the Safety of LLM-Based Search Agents
- CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems
- Tree Search for LLM Agent Reinforcement Learning
- CORE: Full-Path Evaluation of LLM Agents Beyond Final State
- Agent Harness Distillation: Inference-Time Harness Extraction and Exploitation in Autonomous Multi-Agent Systems
- ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents
- A Policy-Driven Runtime Layer for Agentic LLM Serving
- Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
- What Benchmarks Don't Measure: The Case for Evaluating Abstention Competence in Autonomous Agents
- The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
- A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
- MSCoRe: A Benchmark for Multi-Stage Collaborative Reasoning in LLM Agents
- ARE: Scaling Up Agent Environments and Evaluations
- An Evaluation-Centric Paradigm for Scientific Visualization Agents
- AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production
- Who is Introducing the Failure? Automatically Attributing Failures of Multi-Agent Systems via Spectrum Analysis
- Aegis: Automated Error Generation and Attribution for Multi-Agent Systems
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning
- WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research
- WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
- ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization
- Scaling Agents via Continual Pre-training
- Anemoi: A Semi-Centralized Multi-agent System Based on Agent-to-Agent Communication MCP server from Coral Protocol
- Difficulty-Aware Agentic Orchestration for Query-Specific Multi-Agent Workflows
- Abduct, Act, Predict: Scaffolding Causal Inference for Automated Failure Attribution in Multi-Agent Systems
- RAFFLES: Reasoning-based Attribution of Faults for LLM Systems
- Reinforcement Learning Foundations for Deep Research Systems: A Survey
- WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
- SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
- Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects
- AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
- FDABench: A Benchmark for Data Agents on Analytical Queries over Heterogeneous Data
- Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- Physics Supernova: AI Agent Matches Elite Gold Medalists at IPhO 2025
- HiVA: Self-organized Hierarchical Variable Agent via Goal-driven Semantic-Topological Evolution
- How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on τ-bench
- AWorld: Orchestrating the Training Recipe for Agentic AI
- DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
- Foundational Design Principles and Patterns for Building Robust and Adaptive GenAI-Native Systems
- MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
- ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction
- Deep Research: A Survey of Autonomous Research Agents
- Fast, Slow, and Tool-augmented Thinking for LLMs: A Review
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
- MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
- ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
- BrowseMaster: Towards Scalable Web Browsing via Tool-Augmented Programmatic Agent Pair
- WideSearch: Benchmarking Agentic Broad Info-Seeking
- Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL
- SKATE, a Scalable Tournament Eval: Weaker LLMs differentiate between stronger ones using verifiable challenges
- WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging
- ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents
- Characterizing Deep Research: A Benchmark and Formal Definition
- VeriGUI: Verifiable Long-Chain GUI Dataset
- HealthFlow: A Self-Evolving AI Agent with Meta Planning for Autonomous Healthcare Research
- A Survey of LLM-based Deep Search Agents: Paradigm, Optimization, Evaluation, and Challenges
- WebDS: An End-to-End Benchmark for Web-based Data Science
- MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning
- Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training
Discussions
- Meta: Gaia - A Benchmark for General AI Assistants [hn, 36 points, 8 comments]
- haven't seen any of the AI opinion havers on here discussing GAIA yet arxiv.org/abs/2311.12983 [bsky, 1 points, 0 comments]
- GAIA challenges AI with tasks easy for humans but tough for AI, showing a 92% success rate for humans versus 15% for GPT-4. The benchmark includes 466 questions designed to test fundamental abilities [bsky, 0 points, 0 comments]
- tl;dr — AI is dumb. arxiv.org/abs/2311.12983 [bsky, 0 points, 0 comments]
- Pre-print of research paper arxiv.org/abs/2311.12983 [bsky, 0 points, 0 comments]
- Benchmark for General AI Assistants [Mialon+, 2023] GAIA evaluates AI's fundamental abilities, such as web search, image understanding, and file reading, rather than expert knowledge. While humans ach [bsky, 0 points, 0 comments]
Related