The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
2024/08/12 by Chris Lu, Cong Lu, Lu, Chris +10 · 19 voices · 165 citations
Decision Sciences · #Scientific Computing and Data Management
paper · pdf · doi:10.48550/arxiv.2408.06292
Abstract
One of the grand challenges of artificial general intelligence is developing agents capable of conducting scientific research and discovering new knowledge. While frontier models have already been used as aides to human scientists, e.g. for brainstorming ideas, writing code, or prediction tasks, they still conduct only a small part of the scientific process. This paper presents the first comprehensive framework for fully automatic scientific discovery, enabling frontier large language models to perform research independently and communicate their findings. We introduce The AI Scientist, which generates novel research ideas, writes code, executes experiments, visualizes results, describes its findings by writing a full scientific paper, and then runs a simulated review process for evaluation. In principle, this process can be repeated to iteratively develop ideas in an open-ended fashion, acting like the human scientific community. We demonstrate its versatility by applying it to three distinct subfields of machine learning: diffusion modeling, transformer-based language modeling, and learning dynamics. Each idea is implemented and developed into a full paper at a cost of less than 15 per paper. To evaluate the generated papers, we design and validate an automated reviewer, which we show achieves near-human performance in evaluating paper scores. The AI Scientist can produce papers that exceed the acceptance threshold at a top machine learning conference as judged by our automated reviewer. This approach signifies the beginning of a new era in scientific discovery in machine learning: bringing the transformative benefits of AI agents to the entire research process of AI itself, and taking us closer to a world where endless affordable creativity and innovation can be unleashed on the world's most challenging problems. Our code is open-sourced at https://github.com/SakanaAI/AI-Scientist
Cited by
- TCEval: Using Thermal Comfort to Assess Cognitive and Perceptual Abilities of AI
- Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback
- The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research
- Plato-Bio: verification-first biological novelty screening with temporal rediscovery and structural benchmarks
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- CogEEGAgent: Toward Autonomous Cognitive EEG Analysis with Grounded Execution and Selection-Aware Verification
- Agentic Autoresearch for CT Reconstruction
- QMBench: A Research Level Benchmark for Quantum Materials Research
- PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research
- TIB AIssistant: a Platform for AI-Supported Research Across Research Life Cycles
- Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
- Advancing Mathematical Research via Human-AI Interactive Theorem Proving
- The Erosion of LLM Signatures: Can We Still Distinguish Human and LLM-Generated Scientific Ideas After Iterative Paraphrasing?
- DeepCode: Open Agentic Coding
- UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
- AI & Human Co-Improvement for Safer Co-Superintelligence
- Towards an AI Fluid Scientist: LLM-Powered Scientific Discovery in Experimental Fluid Mechanics
- ATHENA: Agentic Team for Hierarchical Evolutionary Numerical Algorithms
- InnoGym: Benchmarking the Innovation Potential of AI Agents
- Measuring Agents in Production
- Simple Agents Outperform Experts in Biomedical Imaging Workflow Optimization
- CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents
- Chain of Unit-Physics: A Primitive-Centric Approach to Scientific Code Synthesis
- Hierarchical AI-Meteorologist: LLM-Agent System for Multi-Scale and Explainable Weather Forecast Reporting
- A Modular LLM-Agent System for Transparent Multi-Parameter Weather Interpretation
- AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
- Cross-Disciplinary Knowledge Retrieval and Synthesis: A Compound AI Architecture for Scientific Discovery
- MirrorMind: Empowering OmniScientist with the Expert Perspectives and Collective Knowledge of Human Scientists
- OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists
- Project Rachel: Can an AI Become a Scholarly Author?
- Accepted with Minor Revisions: Value of AI-Assisted Scientific Writing
- Towards autonomous quantum physics research using LLM agents with access to intelligent tools
- Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
- AlphaResearch: Accelerating New Algorithm Discovery with Language Models
- Scholarly Communications in 2025: An Aerial Evaluation of a System Challenged by <scp>AI</scp> and Much More
- Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author Debates
- AgenticSciML: Collaborative Multi-Agent Systems for Emergent Discovery in Scientific Machine Learning
- Structural Enforcement of Statistical Rigor in AI-Driven Discovery: A Functional Architecture
- Scientific judgment drifts over time in AI ideation
- AgentExpt: Automating AI Experiment Design with LLM-based Resource Retrieval Agent
- Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
- Deep Ideation: Designing LLM Agents to Generate Novel Research Ideas on Scientific Concept Network
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
- Engineering.ai: A Platform for Teams of AI Engineers in Computational Design
- QuantumBench: A Benchmark for Quantum Problem Solving
- Agentic AI: A Comprehensive Survey of Architectures, Applications, and Future Directions
- One Run Is Not an Idea: The Implementation Lottery in Automated Research
- FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents
- Pramana: A Composable, Domain-Specific Backend for Empirical Networking Research
- CREATE: Testing LLMs for Associative Creativity
- AI-based research mentors: Plausible scenarios and ethical issues
- Self-driving laboratories in Japan
- Report of the 2025 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- Playing ZendoWorld: Challenging AI Agents on Active Visual Concept Induction
- Autodata: An agentic data scientist to create high quality synthetic data
- The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science
- The rise of the research automaton: science as process or product in the era of generative AI?
- Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft
- AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
- Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis
- Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts
- Evidence-Bound Autonomous Research (EviBound): A Governance Framework for Eliminating False Claims
- A Survey of AI Scientists
- HRM-Agent: Training a recurrent reasoning model in dynamic environments using reinforcement learning
- How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
- AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing
- Magellan: Guided MCTS for Latent Space Exploration and Novelty Generation
- Shoot First, Ask Questions Later? Building Rational Agents that Explore and Act Like People
- Co-Designing Quantum Codes with Transversal Diagonal Gates via Multi-Agent Systems
- Build Your Personalized Research Group: A Multiagent Framework for Continual and Interactive Science Automation
- ResearchGPT: Benchmarking and Training LLMs for End-to-End Computer Science Research Workflows
- Machine Text Detectors are Membership Inference Attacks
- Cultural Alien Sampler: Open-ended art generation balancing originality and coherence
- BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
- AcademicEval: Live Long-Context LLM Benchmark
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- From AutoRecSys to AutoRecLab: A Call to Build, Evaluate, and Govern Autonomous Recommender-Systems Research Labs
- Foundation Models for Scientific Discovery: From Paradigm Enhancement to Paradigm Transition
- Agentic Discovery: Closing the Loop with Cooperative Agents
- CiteGuard: Faithful Citation Attribution for LLMs via Retrieval-Augmented Validation
- Spec-Driven AI for Science: The ARIA Framework for Automated and Reproducible Data Analysis
- FML-bench: A Benchmark for Automatic ML Research Agents Highlighting the Importance of Exploration Breadth
- Scheming Ability in LLM-to-LLM Strategic Interactions
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
- SIA: Self Improving AI with Harness & Weight Updates
- The Last Human-Written Paper: Agent-Native Research Artifacts
- An Alternative Trajectory for Generative AI
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- AutoPR: Let's Automate Your Academic Promotion!
- Maple: A Multi-agent System for Portable Deep Learning across Clusters
- FlowSearch: Advancing deep research with dynamic structured knowledge flow
- Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
- From Data to Rewards: a Bilevel Optimization Perspective on Maximum Likelihood Estimation
- Hypothesis Hunting with Evolving Networks of Autonomous Scientific Agents
- Evolving and Executing Research Plans via Double-Loop Multi-Agent Collaboration
- Inefficiencies of Meta Agents for Agent Design
- TinyScientist: An Interactive, Extensible, and Controllable Framework for Building Research Agents
- Scientific Algorithm Discovery by Augmenting AlphaEvolve with Deep Research
- RareAgent: Self-Evolving Reasoning for Drug Repurposing in Rare Diseases
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- Fusing Multi- and Hyperspectral Satellite Data for Harmful Algal Bloom Monitoring with Self-Supervised and Hierarchical Deep Learning
- Self-Improvement in Multimodal Large Language Models: A Survey
- JoyAgent-JDGenie: Technical Report on the GAIA
- DeepScientist: Advancing Frontier-Pushing Scientific Findings Progressively
- AutoLabs: Cognitive Multi-Agent Systems with Self-Correction for Autonomous Chemical Experimentation
- Scaling Generalist Data-Analytic Agents
- Agentic Exploration of Physics Models
- Memory Transfer Planning: LLM-driven Context-Aware Code Adaptation for Robot Manipulation
- NAIPv2: Debiased Pairwise Learning for Efficient Paper Quality Estimation
- CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems
- PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
- Beyond efficiency: How artificial intelligence (AI) will reshape scientific inquiry and the publication process
- Guiding Evolution of Artificial Life Using Vision-Language Models
- FormalML: A Benchmark for Evaluating Formal Subgoal Completion in Machine Learning Theory
- MotivGraph-SoIQ: Integrating Motivational Knowledge Graphs and Socratic Dialogue for Enhanced LLM Ideation
- Bridging the Gap Between Scientific Laws Derived by AI Systems and Canonical Knowledge via Abductive Inference with AI-Noether
- Drawing out What They Struggle to say: AI-Augmented Analysis of Projective Techniques in Qualitative Health Research
- TCA-SIR: Learning Target-Conditioned Abstractions for Scientific Inspiration Retrieval
- Baikal: Structured Search for Deep Research over Data Lakes
- Rehearse: Stepping Back from the Confidence Cliff in Self-Improving Autoresearch
- MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck
- CLVisc Agent for autonomous relativistic hydrodynamics studies
- Adsorb-Agent: autonomous identification of stable adsorption configurations <i>via</i> a large language model agent
- Inspectable AI for Science: A Research Object Approach to Generative AI Governance
- Spacer: Towards Engineered Scientific Inspiration
- Tenure Under Pressure: Simulating the Disruptive Effects of AI on Academic Publishing
- OpenLens AI: Fully Autonomous Research Agent for Health Infomatics
- Mimicking the Physicist's Eye:A VLM-centric Approach for Physics Formula Discovery
- Automated Generation of Research Workflows from Academic Papers: A Full-text Mining Framework
- Evalet: Evaluating Large Language Models through Functional Fragmentation
- Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
- How are Scientific Concepts Birthed? Typing Rules of Concept Formation in Theoretical Physics Reasoning
- Solve it with EASE
- The Need for Verification in AI-Driven Scientific Discovery
- From Tape Reels to Global Access: A History and Future Vision of NASA's Scientific Data Management
- Charting the Future of Scholarly Knowledge with AI: A Community Perspective
- Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning
- The Statistical Validation of Innovation Lens
- AI Agents for Photonic Integrated Circuit Design Automation
- Deep Research: A Survey of Autonomous Research Agents
- AgentCDM: Enhancing Multi-Agent Collaborative Decision-Making via ACH-Inspired Structured Reasoning
- A Comprehensive Review of AI Agents: Transforming Possibilities in Technology and Beyond
- A Multi-Task Evaluation of LLMs' Processing of Academic Text Input
- SGSimEval: A Comprehensive Multifaceted and Similarity-Enhanced Benchmark for Automatic Survey Generation Systems
- Beyond "Not Novel Enough": Enriching Scholarly Critique with LLM-Assisted Feedback
- Considering the ethics of large machine learning models in the chemical sciences
- ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
- ReviewRL: Towards Automated Scientific Review with RL
- What are the limits to biomedical research acceleration through general-purpose AI?
- MiGrATe: Mixed-Policy GRPO for Adaptation at Test-Time
- InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling
- Multi-agent systems for chemical engineering: A review and perspective
- Deep learning methods for 2D material electronic properties
- Navigating Through Paper Flood: Advancing LLM-based Paper Evaluation through Domain-Aware Retrieval and Latent Reasoning
- Beyond Brainstorming: What Drives High-Quality Scientific Ideas? Lessons from Multi-Agent Collaboration
- An Auditable Agent Platform For Automated Molecular Optimisation
- Autonomous Inorganic Materials Discovery via Multi-Agent Physics-Aware Scientific Reasoning
- Cognitive Loop via In-Situ Optimization: Self-Adaptive Reasoning for Science
- A Survey on AgentOps: Categorization, Challenges, and Future Directions
- Bayes-Entropy Collaborative Driven Agents for Research Hypotheses Generation and Optimization
- MaRGen: Multi-Agent LLM Approach for Self-Directed Market Research and Analysis
- BioDisco: Multi-agent hypothesis generation with dual-mode evidence, iterative feedback and temporal evaluation
- SynAdapt: Learning Adaptive Reasoning in Large Language Models via Synthetic Continuous Chain-of-Thought
- How Far Are AI Scientists from Changing the World?
Discussions
- Don’t share the revulsion some feel about this paper, but I do think it’s flawed. Their automated evaluator is validated on human papers (where it will learn to detect a certain set of problems) and t [bsky, 6 points, 2 comments]
- An AI research assistant that was faced with a strict time limit tried to rewrite the code to remove the time limit, instead of completing the task. arxiv.org/pdf/2408.06292 [bsky, 5 points, 2 comments]
- Humans were a mistake. We should've stopped with fire. [bsky, 3 points, 0 comments]
- A paper that dares to ask, "how can we make the replication crisis even worse? How about adding LLM hallucinations to the mix?"
Never mind the likely plagiarism that will inevitably result from spitti [bsky, 2 points, 1 comments]
- The AI Scientist: Towards Automated Open-Ended Scientific Discovery [hn, 2 points, 0 comments]
- There are people that would get AI to write part or all of the text of their papers. I've even been to talks where people are excited about when we can not do experiments and just think about question [bsky, 2 points, 1 comments]
- This is fine. arxiv.org/abs/2408.06292 (Mein Link für heute) #AIScientist #AutonomeSoftware [bsky, 2 points, 0 comments]
- THIS HAS A PREPRINT FML [bsky, 1 points, 0 comments]
- bro literally what is the matter with you
https://arxiv.org/abs/2408.06292 [bsky, 1 points, 0 comments]
- The AI Scientist: Towards Automated Open-Ended Scientific Discovery [hn, 1 points, 0 comments]
- The AI Scientist: Towards Automated Open-Ended Scientific Discovery [hn, 1 points, 0 comments]
- Kauempaa katsoen tiede ja tutkimus voivat näyttää uusien tulosten tuottamiselta. Silloin voi perustellusti kysyä, eikö tekoäly voisi tehdä sitä? Ei tuloksia hatarasti arvaten kuten kielimallit, vaan j [bsky, 1 points, 1 comments]
- Sweet! Robots have come for my job (robotics research): arxiv.org/pdf/2408.06292 [bsky, 0 points, 1 comments]
- arxiv.org/abs/2408.06292 [bsky, 0 points, 0 comments]
- Finally, an academics wet dream has come true: end-to-end paper generation. doi.org/10.48550/arX... [bsky, 0 points, 0 comments]
- This new preprint might worry our CS colleagues.
"We introduce The AI Scientist, which generates novel research ideas, writes code, executes experiments, visualizes results, describes its findings by [bsky, 0 points, 1 comments]
- Oh wonderful - looking forward to more ‘frameworks’ for polluting journals with LLM-generated slop. arxiv.org/abs/2408.06292 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2408.06292
MDPI is going to be delighted. Best case this will force us to reconsider how we measure scientific productivity. Using LLMs to render the current working mode of academia ad [bsky, 0 points, 2 comments]
- arxiv.org/abs/2408.06292 [bsky, 0 points, 0 comments]
Related