Accelerating scientific discovery with Co-Scientist
2025/02/26 by Juraj Gottweis, Gottweis, Juraj, Wei‐Hung Weng +97 · 8 voices · 97 citations
Decision Sciences · Biochemistry, Genetics and Molecular Biology · #Academic Publishing and Open Access #Genetics, Bioinformatics, and Biomedical Research #Scientific Computing and Data Management
paper · pdf · doi:10.1038/s41586-026-10644-y
Abstract
Scientific discovery is driven by scientists generating novel hypotheses for complex problems that undergo rigorous experimental validation. To augment this process, we introduce Co-Scientist, a multi-agent AI system built on Gemini for structured scientific thinking and hypothesis generation. Co-Scientist aims to help scientists discover new original knowledge. Conditioned on their research objectives and prior scientific evidence, it formulates demonstrably novel research hypotheses for experimental verification. The system’s design involves agents continuously generating, critiquing and refining hypotheses accelerated by scaling test-time compute. Key contributions include: (1) a multi-agent architecture with an asynchronous task execution framework for flexible compute scaling; (2) a tournament evolution process for self-improving hypotheses generation. Automated evaluations show continued benefits of test-time compute scaling, improving hypothesis quality over time. While general purpose, we focus the validation in three biomedical applications: drug repurposing, novel target discovery 1, and explaining mechanisms of anti-microbial resistance 2. Specifically, Co-Scientist helped identify new drug repurposing candidates and synergistic combination therapies for acute myeloid leukemia, which were validated through in vitro experiments. These real-world validations demonstrate the potential of Co-Scientist to accelerate scientific discovery and usher in an era of AI empowered scientists.
Citations
Cited by
- Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity
- IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation
- WARA: A Closed-Loop Multi-Agent Framework for Wireless Optimization Autoresearch
- IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation
- AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning
- Agents in the Wild: Where Research Meets Deployment
- LLM-Driven Cross-Paradigm Design for Quantum Optimal Control
- An Intelligent Infrastructure as a Foundation for Modern Science
- Schema-Bound LLM Control of Scientific Instrumentation through Model Context Protocol Skills
- BrainPilot: Automating Brain Discovery with Agentic Research
- SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery
- AutoSynthesis: An agentic system for automated meta-analysis
- Autonomous mechanistic discovery of colorectal cancer vulnerabilities via multi-scale AI swarms
- AI Research Agents Narrow Scientific Exploration
- MechAInistic: An LLM-guided Multi-Agent System for Reasoning over Genome-Scale Constraint-Based Metabolic Models
- AI co-mathematician: Accelerating mathematicians with agentic AI
- Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
- Paper2Agent: Reimagining Research Papers As Interactive and Reliable AI Agents
- Failures Reveal What Metrics Miss: An Evidence-Driven Agent for Recursive Refinement of ECG Classifiers
- When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
- How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?
- Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
- PaperBanana: Automating Academic Illustration for AI Scientists
- Agentic AI-enabled discovery across large-scale sleep physiology
- Agentic AI for Scientific Reasoning in Autonomous Quantum Sensing Experiments
- Extremal Chowla sets and their linear analogues: A human-AI mathematical investigation using Co-Scientist
- QMBench: A Research Level Benchmark for Quantum Materials Research
- FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs
- Bohrium + SciMaster: Building the Infrastructure and Ecosystem for Agentic Science at Scale
- PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research
- Multimodal LLMs for Historical Dataset Construction from Archival Image Scans: German Patents (1877-1918)
- ReX-MLE: The Autonomous Agent Benchmark for Medical Imaging Challenges
- Estimating problem difficulty without ground truth using Large Language Model comparisons
- Towards a Science of Scaling Agent Systems
- AI & Human Co-Improvement for Safer Co-Superintelligence
- auto-psych: Automating the science of mind using agent-driven theory discovery and experimentation
- Towards an AI Fluid Scientist: LLM-Powered Scientific Discovery in Experimental Fluid Mechanics
- InnoGym: Benchmarking the Innovation Potential of AI Agents
- E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing
- SelfAI: Building a Self-Training AI System with LLM Agents
- Cross-Disciplinary Knowledge Retrieval and Synthesis: A Compound AI Architecture for Scientific Discovery
- OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists
- ATLAS: A High-Difficulty, Multidisciplinary Benchmark for Frontier Scientific Reasoning
- Towards autonomous quantum physics research using LLM agents with access to intelligent tools
- Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
- Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
- Advancing Subsurface Discovery and Geothermal Monitoring with an Agentic Artificial Intelligence Framework
- A Hierarchical Multi-Agent System for Autonomous Discovery in Geoscientific Data Archives
- Deep Ideation: Designing LLM Agents to Generate Novel Research Ideas on Scientific Concept Network
- Engineering.ai: A Platform for Teams of AI Engineers in Computational Design
- PUDA: An AI-Native Hardware Harness for Self-Driving Laboratories
- AI-guided discovery of atypical protein assemblies
- AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation
- AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
- CooperBench: Why Coding Agents Cannot be Your Teammates Yet
- A Survey of AI Scientists
- AutoSciDACT: Automated Scientific Discovery through Contrastive Embedding and Hypothesis Testing
- Build Your Personalized Research Group: A Multiagent Framework for Continual and Interactive Science Automation
- From AutoRecSys to AutoRecLab: A Call to Build, Evaluate, and Govern Autonomous Recommender-Systems Research Labs
- Helmsman: Autonomous Synthesis of Federated Learning Systems via Collaborative LLM Agents
- Ax-Prover: A Deep Reasoning Agentic Framework for Theorem Proving in Mathematics and Quantum Physics
- The Agentic Garden of Forking Paths
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
- Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
- Maple: A Multi-agent System for Portable Deep Learning across Clusters
- InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation
- Hypothesis Hunting with Evolving Networks of Autonomous Scientific Agents
- Scientific Algorithm Discovery by Augmenting AlphaEvolve with Deep Research
- Large Language Models Achieve Gold Medal Performance at the International Olympiad on Astronomy & Astrophysics (IOAA)
- Multimodal AI agents for capturing and sharing laboratory practice
- AgentHub: A Research Agenda for Agent Sharing Infrastructure
- DeepScientist: Advancing Frontier-Pushing Scientific Findings Progressively
- FormalML: A Benchmark for Evaluating Formal Subgoal Completion in Machine Learning Theory
- Evaluating Agentic Bioinformatics through Function, Evidence, and Validation
- EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents
- From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery
- Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology
- Reinforced Generation of Combinatorial Structures: Hardness of Approximation
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- Co-Investigator AI: The Rise of Agentic AI for Smarter, Trustworthy AML Compliance Narratives
- Solve it with EASE
- OwkinZero: Accelerating Biological Discovery with AI
- The Need for Verification in AI-Driven Scientific Discovery
- What Are Research Hypotheses?
- Charting the Future of Scholarly Knowledge with AI: A Community Perspective
- AI Agents for Photonic Integrated Circuit Design Automation
- LARC: Towards Human-level Constrained Retrosynthesis Planning through an Agentic Framework
- ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
- What are the limits to biomedical research acceleration through general-purpose AI?
- InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling
- BrowseMaster: Towards Scalable Web Browsing via Tool-Augmented Programmatic Agent Pair
- Multi-agent systems for chemical engineering: A review and perspective
- An Auditable Agent Platform For Automated Molecular Optimisation
- A Multi-Agent System for Complex Reasoning in Radiology Visual Question Answering
- Cognitive Loop via In-Situ Optimization: Self-Adaptive Reasoning for Science
- SimuRA: A World-Model-Driven Simulative Reasoning Architecture for General Goal-Oriented Agents
- How Far Are AI Scientists from Changing the World?
Discussions
- Towards an AI Co-Scientist [hn, 47 points, 17 comments]
- An AI Co-Scientist for Hypothesis Generation from Google DeepMind [hn, 4 points, 0 comments]
- “Toward an AI co-scientist”: arxiv.org/abs/2502.18864 [bsky, 1 points, 1 comments]
- Another paper that I should read. I read an earlier paper from this group that was very scattered (too many topics) and unsupported. Hoping this paper steps up the rigor. doi.org/10.1038/s415... [bsky, 1 points, 0 comments]
- In this paper, they demonstrate an AI that will propose hypotheses and select the best one via an election process arxiv.org/abs/2502.18864 new hyptheses are novel [bsky, 0 points, 0 comments]
- Towards an AI Co-Scientist https://arxiv.org/abs/2502.18864 (https://news.ycombinator.com/item?id=43205755) [bsky, 0 points, 0 comments]
- Welcome to AI Co-Scientist, a #Gemini 2.0-based system which can generate, debate, and evolve approach to hypothesis generation, inspired by the scientific method and accelerated by scaling test-time [bsky, 0 points, 0 comments]
- Towards an AI Co-Scientist https://arxiv.org/abs/2502.18864 [bsky, 0 points, 0 comments]
Related