The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
2025/04/10 by Yamada, Yutaro, Lange, Robert Tjarko, Lu, Cong +5 · 69 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
paper · doi:10.48550/arxiv.2504.08066
Abstract
AI is increasingly playing a pivotal role in transforming how scientific discoveries are made. We introduce The AI Scientist-v2, an end-to-end agentic system capable of producing the first entirely AI generated peer-review-accepted workshop paper. This system iteratively formulates scientific hypotheses, designs and executes experiments, analyzes and visualizes data, and autonomously authors scientific manuscripts. Compared to its predecessor (v1, Lu et al., 2024 arXiv:2408.06292), The AI Scientist-v2 eliminates the reliance on human-authored code templates, generalizes effectively across diverse machine learning domains, and leverages a novel progressive agentic tree-search methodology managed by a dedicated experiment manager agent. Additionally, we enhance the AI reviewer component by integrating a Vision-Language Model (VLM) feedback loop for iterative refinement of content and aesthetics of the figures. We evaluated The AI Scientist-v2 by submitting three fully autonomous manuscripts to a peer-reviewed ICLR workshop. Notably, one manuscript achieved high enough scores to exceed the average human acceptance threshold, marking the first instance of a fully AI-generated paper successfully navigating a peer review. This accomplishment highlights the growing capability of AI in conducting all aspects of scientific research. We anticipate that further advancements in autonomous scientific discovery technologies will profoundly impact human knowledge generation, enabling unprecedented scalability in research productivity and significantly accelerating scientific breakthroughs, greatly benefiting society at large. We have open-sourced the code at https://github.com/SakanaAI/AI-Scientist-v2 to foster the future development of this transformative technology. We also discuss the role of AI in science, including AI safety.
Cited by
- Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse
- Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis
- Bridging the Gap on AI-Assisted Scientific Software Development Through Transparency and Traceability
- Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact
- PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research
- Towards an AI Fluid Scientist: LLM-Powered Scientific Discovery in Experimental Fluid Mechanics
- ARCADIA: Scalable Causal Discovery for Corporate Bankruptcy Analysis Using Agentic AI
- Beyond Protein Language Models: An Agentic LLM Framework for Mechanistic Enzyme Design
- MirrorMind: Empowering OmniScientist with the Expert Perspectives and Collective Knowledge of Human Scientists
- OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- TacEleven: generative tactic discovery for football open play
- Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
- Mimosa Framework: Toward Evolving Multi-Agent Systems for Scientific Research
- Structural Enforcement of Statistical Rigor in AI-Driven Discovery: A Functional Architecture
- Preserving security in a world with powerful AI Considerations for the future Defense Architecture
- Scientific judgment drifts over time in AI ideation
- AgentExpt: Automating AI Experiment Design with LLM-based Resource Retrieval Agent
- Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
- A Hierarchical Multi-Agent System for Autonomous Discovery in Geoscientific Data Archives
- Deep Ideation: Designing LLM Agents to Generate Novel Research Ideas on Scientific Concept Network
- Engineering.ai: A Platform for Teams of AI Engineers in Computational Design
- QuantumBench: A Benchmark for Quantum Problem Solving
- FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents
- Pramana: A Composable, Domain-Specific Backend for Empirical Networking Research
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- SynthWorlds: Controlled Parallel Worlds for Disentangling Reasoning and Knowledge in Language Models
- A Survey of AI Scientists
- AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing
- Build Your Personalized Research Group: A Multiagent Framework for Continual and Interactive Science Automation
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- From AutoRecSys to AutoRecLab: A Call to Build, Evaluate, and Govern Autonomous Recommender-Systems Research Labs
- Foundation Models for Scientific Discovery: From Paradigm Enhancement to Paradigm Transition
- CiteGuard: Faithful Citation Attribution for LLMs via Retrieval-Augmented Validation
- Ax-Prover: A Deep Reasoning Agentic Framework for Theorem Proving in Mathematics and Quantum Physics
- FML-bench: A Benchmark for Automatic ML Research Agents Highlighting the Importance of Exploration Breadth
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
- The Last Human-Written Paper: Agent-Native Research Artifacts
- Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
- IoDResearch: Deep Research on Private Heterogeneous Data via the Internet of Data
- Evolving and Executing Research Plans via Double-Loop Multi-Agent Collaboration
- Inefficiencies of Meta Agents for Agent Design
- TinyScientist: An Interactive, Extensible, and Controllable Framework for Building Research Agents
- RareAgent: Self-Evolving Reasoning for Drug Repurposing in Rare Diseases
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- Foundation models for equation discovery in high energy physics
- Self-Improvement in Multimodal Large Language Models: A Survey
- DeepScientist: Advancing Frontier-Pushing Scientific Findings Progressively
- InfiAgent: Self-Evolving Pyramid Agent Framework for Infinite Scenarios
- FormalML: A Benchmark for Evaluating Formal Subgoal Completion in Machine Learning Theory
- MotivGraph-SoIQ: Integrating Motivational Knowledge Graphs and Socratic Dialogue for Enhanced LLM Ideation
- Rehearse: Stepping Back from the Confidence Cliff in Self-Improving Autoresearch
- CLVisc Agent for autonomous relativistic hydrodynamics studies
- Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
- Spacer: Towards Engineered Scientific Inspiration
- Tenure Under Pressure: Simulating the Disruptive Effects of AI on Academic Publishing
- The Need for Verification in AI-Driven Scientific Discovery
- Deep Research: A Survey of Autonomous Research Agents
- ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
- Multi-agent systems for chemical engineering: A review and perspective
- Bayes-Entropy Collaborative Driven Agents for Research Hypotheses Generation and Optimization
- How Far Are AI Scientists from Changing the World?
- Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts
- Grounded autonomous research: a fault-tolerant LLM pipeline from corpus to manuscript in frontier computational physics
- Towards Execution-Grounded Automated AI Research
Related