Data Interpreter: An LLM Agent For Data Science
2024/02/28 by Sirui Hong, Yizhang Lin, Hong, Sirui +53 · 3 voices · 81 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer science #Data science #FOS: Computer and information sciences #Interpreter #Machine Learning (cs.LG) #Programming language #Semantic Web and Ontologies #cs.AI #cs.LG
paper · pdf · doi:10.48550/arxiv.2402.18679
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/02/28 · arxiv published 2024/02/28 · arxiv updated 2024/10/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Large Language Model (LLM)-based agents have shown effectiveness across many applications. However, their use in data science scenarios requiring solving long-term interconnected tasks, dynamic data adjustments and domain expertise remains challenging. Previous approaches primarily focus on individual tasks, making it difficult to assess the complete data science workflow. Moreover, they struggle to handle real-time changes in intermediate data and fail to adapt dynamically to evolving task dependencies inherent to data science problems. In this paper, we present Data Interpreter, an LLM-based agent designed to automatically solve various data science problems end-to-end. Our Data Interpreter incorporates two key modules: 1) Hierarchical Graph Modeling, which breaks down complex problems into manageable subproblems, enabling dynamic node generation and graph optimization; and 2) Programmable Node Generation, a technique that refines and verifies each subproblem to iteratively improve code generation results and robustness. Extensive experiments consistently demonstrate the superiority of Data Interpreter. On InfiAgent-DABench, it achieves a 25% performance boost, raising accuracy from 75.9% to 94.9%. For machine learning and open-ended tasks, it improves performance from 88% to 95%, and from 60% to 97%, respectively. Moreover, on the MATH dataset, Data Interpreter achieves remarkable performance with a 26% improvement compared to state-of-the-art baselines. The code is available at https://github.com/geekan/MetaGPT.
Cited by
- Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse
- SCTA: An Agentic Framework for Stable and Interpretable Target Gene Discovery from Single-Cell RNA Sequencing
- SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents
- Can Agentic AI Match the Performance of Human Data Scientists?
- Can AI autonomously build, operate, and use the entire data stack?
- DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
- An Empirical Study of Agent Developer Practices in AI Agent Frameworks
- Toward Automated and Trustworthy Scientific Analysis and Visualization with LLM-Generated Code
- DataSage: Multi-agent Collaboration for Insight Discovery with External Knowledge Retrieval, Multi-role Debating, and Multi-path Reasoning
- UniDataBench: Evaluating Data Analytics Agents Across Structured and Unstructured Data
- A Hierarchical Multi-Agent System for Autonomous Discovery in Geoscientific Data Archives
- Optimal-Agent-Selection: State-Aware Routing Framework for Efficient Multi-Agent Collaboration
- What's the next frontier for Data-centric AI? Data Savvy Agents
- Data Analysis and Performance Evaluation of Simulation Deduction Based on LLMs
- UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks
- Autodata: An agentic data scientist to create high quality synthetic data
- VDSAgents: A PCS-Guided Multi-Agent System for Veridical Data Science Automation
- A Survey of Data Agents: Emerging Paradigm or Overstated Hype?
- DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
- InferA: A Smart Assistant for Cosmological Ensemble Data
- AwareCompiler: Agentic Context-Aware Compiler Optimization via a Synergistic Knowledge-Data Driven Framework
- Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
- ConDABench: Interactive Evaluation of Language Models for Data Analysis
- LLM-FS-Agent: A Deliberative Role-based Large Language Model Architecture for Transparent Feature Selection
- Trade in Minutes! Rationality-Driven Agentic System for Quantitative Financial Trading
- SPOGW: a Score-based Preference Optimization method via Group-Wise comparison for workflows
- DyFlow: Dynamic Workflow Framework for Agentic Reasoning
- LLM-Based Multi-Agent Blackboard System for Information Discovery in Data Science
- Scaling Generalist Data-Analytic Agents
- Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence
- RobustFlow: Towards Robust Agentic Workflow Generation
- A State-of-the-Art SQL Reasoning Model using RLVR
- Difficulty-Aware Agentic Orchestration for Query-Specific Multi-Agent Workflows
- Jupiter: Enhancing LLM Data Analysis Capabilities via Notebook and Inference-Time Value-Guided Search
- AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
- A Survey on Open Dataset Search in the LLM Era: Retrospectives and Perspectives
- Datarus-R1: An Adaptive Multi-Step Reasoning LLM for Automated Data Analysis
- Tabularis Formatus: Predictive Formatting for Tables
- Non-programmers Assessing AI-Generated Code: A Case Study of Business Users Analyzing Data
- Polymath: A Self-Optimizing Agent with Dynamic Hierarchical Workflow
- Large Language Model-based Data Science Agent: A Survey
- WebDS: An End-to-End Benchmark for Web-based Data Science
- GPT-4.1 Sets the Standard in Automated Experiment Design Using Novel Python Libraries
- Large Language Model Agent for Structural Drawing Generation Using ReAct Prompt Engineering and Retrieval Augmented Generation
- Data Agent: A Holistic Architecture for Orchestrating Data+AI Ecosystems
- G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems
- CoMind: Towards Community-Driven Agents for Machine Learning Engineering
- ReCode: Updating Code API Knowledge with Reinforcement Learning
- Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study
- Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim → Evidence Reasoning
- AutoMind: Adaptive Knowledgeable Agent for Automated Data Science
- AgentSwift: Efficient LLM Agent Design via Value-guided Hierarchical Search
- Optimization Problem Solving Can Transition to Evolutionary Agentic Workflows
- Can AI Master Econometrics? Evidence from Econometrics AI Agent on Expert-Level Tasks
- Knowledge Augmented Complex Problem Solving with Large Language Models: A Survey
- SimuGen: Multi-modal Agentic Framework for Constructing Block Diagram-Based Simulation Models
- MLE-STAR: Machine Learning Engineering Agent via Search and Targeted Refinement
- LLM-Agent-Controller: A Universal Multi-Agent Large Language Model System as a Control Engineer
- AssistedDS: Benchmarking How External Domain Knowledge Assists LLMs in Automated Data Science
- GABench: A Comprehensive Benchmark for Evaluating LLM Agents on Graph Analysis Tasks
- Beyond Solution-Centric Search: Adaptive Inquiry and Knowledge Revision for Autonomous ML Engineering
- Large Language Models for Predictive Analysis: How Far Are They?
- Adaptive Plan-Execute Framework for Smart Contract Security Auditing
- DSMentor: Enhancing Data Science Agents with Curriculum Learning and Online Knowledge Accumulation
- Simulation Agent: A Framework for Integrating Simulation and Large Language Models for Enhanced Decision-Making
- Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems
- MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering
- AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems
- LLMs can construct powerful representations and streamline sample-efficient supervised learning
- Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models
- "When to Hand Off, When to Work Together": Expanding Human-Agent Co-Creative Collaboration through Concurrent Interaction
- DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science
- From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
- Data Agents: Levels, State of the Art, and Open Problems
- Keyword search is all you need: Achieving RAG-Level Performance without vector databases using agentic tool use
- VERITAS: A Multi-Agent Co-Scientist for Verifiable Image-Derived Hypothesis Testing
- AgentAda: Skill-Adaptive Data Analytics for Tailored Insight Discovery
- SciSciGPT: Advancing Human-AI Collaboration in the Science of Science
- ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines
- Weak-for-Strong: Training Weak Meta-Agent to Harness Strong Executors
- Agentic Knowledgeable Self-awareness
Discussions
Related