Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
2025/10/10 by Zhou, Lianhao, Ling, Hongyi, Fu, Cong +14 · 3 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2510.09901
Abstract
Computing has long served as a cornerstone of scientific discovery. Recently, a paradigm shift has emerged with the rise of large language models (LLMs), introducing autonomous systems, referred to as agents, that accelerate discovery across varying levels of autonomy. These language agents provide a flexible and versatile framework that orchestrates interactions with human scientists, natural language, computer language and code, and physics. This paper presents our view and vision of LLM-based scientific agents and their growing role in transforming the scientific discovery lifecycle, from hypothesis discovery, experimental design and execution, to result analysis and refinement. We critically examine current methodologies, emphasizing key innovations, practical achievements, and outstanding limitations. Additionally, we identify open research challenges and outline promising directions for building more robust, generalizable, and adaptive scientific agents. Our analysis highlights the transformative potential of autonomous agents to accelerate scientific discovery across diverse domains.
Citations
- Democratizing AI scientists using ToolUniverse
- ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution
- From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery
- LARC: Towards Human-level Constrained Retrosynthesis Planning through an Agentic Framework
- FROGENT: An End-to-End Full-process Drug Design Agent
- WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent
- LLM-based Multi-Agent Copilot for Quantum Sensor
- Beyond Brainstorming: What Drives High-Quality Scientific Ideas? Lessons from Multi-Agent Collaboration
- Bayes-Entropy Collaborative Driven Agents for Research Hypotheses Generation and Optimization
- SciToolAgent: A Knowledge Graph-Driven Scientific Agent for Multi-Tool Integration
- Language Models for Controllable DNA Sequence Design
- AlphaGenome: advancing regulatory variant effect prediction with a unified DNA sequence model
- Iterative Distillation for Reward-Guided Fine-Tuning of Diffusion Models in Biomolecular Design
- PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier
- Efficient Prediction of SO(3)-Equivariant Hamiltonian Matrices via SO(2) Local Frames
- Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
- Training a Scientific Reasoning Model for Chemistry
- Exploiting LLMs for Automatic Hypothesis Assessment via a Logit-Based Calibrated Prior
- Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
- Foundation Molecular Grammar: Multi-Modal Foundation Models Induce Interpretable Molecular Graph Languages
- From Reasoning to Learning: A Survey on Hypothesis Discovery and Rule Learning with Large Language Models
- WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
- Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models
- Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
- From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery
- CodePDE: An Inference Framework for LLM-driven PDE Solver Generation
- El Agente: An autonomous agent for quantum chemistry
- RM-R1: Reward Modeling as Reasoning
- WebThinker: Empowering Large Reasoning Models with Deep Research Capability
- Confidence in Large Language Model Evaluation: A Bayesian Approach to Limited-Sample Challenges
- From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
- Sparks: Multi-Agent Artificial Intelligence Model Discovers Protein Design Principles
- From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs
- HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
- A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models
- Miiyuu/scAgent: Universal Cell Type Annotation via a LLM
- ChemToolAgent: The Impact of Tools on Language Agents for Chemistry Problem Solving
- Towards Scientific Intelligence: A Survey of LLM-based Scientific Agents
- AstroAgents: A Multi-Agent AI for Hypothesis Generation from Mass Spectrometry Data
- PharmAgents: Building a Virtual Pharma with Large Language Model Agents
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
- TxAgent: An AI Agent for Therapeutic Reasoning Across a Universe of Tools
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Agentic AI for Scientific Discovery: A Survey of Progress, Challenges, and Future Directions
- Dynamic Search for Inference-Time Alignment in Diffusion Models
- Invariant Tokenization of Crystalline Materials for Language Model Enabled Generation
- Accelerating scientific discovery with Co-Scientist
- RAG-Enhanced Collaborative LLM Agents for Drug Discovery
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- Reward-Guided Iterative Refinement in Diffusion Models at Test-Time with Applications to Protein and DNA Design
- LIDDIA: Language-based Intelligent Drug Discovery Agent
- Learning to Discover Regulatory Elements for Gene Expression Prediction
- LLM Agents Making Agent Tools
- A-MEM: Agentic Memory for LLM Agents
- PlanGenLLMs: A Modern Survey of LLM Planning Capabilities
- Agentic End-to-End De Novo Protein Design for Tailored Dynamics Using a Language Diffusion Model
- Automated Hypothesis Validation with Agentic Sequential Falsifications
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- Hypothesis Generation for Materials Discovery and Design Using Goal-Driven and Constraint-Guided LLM Agents
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- FRAG: A Flexible Modular Framework for Retrieval-Augmented Generation based on Knowledge Graphs
- Inference-Time Alignment in Diffusion Models with Reward-Guided Generation: Tutorial and Review
- ChemAgent: Self-updating Library in Large Language Models Improves Chemical Reasoning
- BioAgents: Democratizing Bioinformatics Analysis with Multi-Agent Systems
- OpenFOAMGPT: a RAG-Augmented LLM Agent for OpenFOAM-Based Computational Fluid Dynamics
- Large Physics Models: Towards a collaborative approach with Large Language Models and Foundation Models
- LLM4SR: A Survey on Large Language Models for Scientific Research
- A Retrieval-Augmented Knowledge Mining Method with Deep Thinking LLMs for Biomedical Research and Clinical Support
- Towards Scientific Discovery with Generative AI: Progress, Opportunities, and Challenges
- From Intention To Implementation: Automating Biomedical Research via LLMs
- Agents for self-driving laboratories applied to quantum computing
- A Review on Scientific Knowledge Extraction using Large Language Models in Biomedical Sciences
- A Multi-agent Framework for Physical Laws Discovery
- DrugAgent: Automating AI-aided Drug Discovery Programming through LLM Multi-Agent Collaboration
- Physics-Informed Autonomous LLM Agents for Explainable Power Electronics Modulation Design
- MatPilot: an LLM-enabled AI Materials Scientist under the Framework of Human-Machine Collaboration
- AutoProteinEngine: A Large Language Model Driven Agent Framework for Multimodal AutoML in Protein Engineering
- GIS Copilot: Towards an Autonomous GIS Agent for Spatial Analysis
- Improving Scientific Hypothesis Generation with Knowledge Grounded Large Language Models
- MatViX: Multimodal Information Extraction from Visually Rich Articles
- GPT-4o System Card
- Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System
- Agent S: An Open Agentic Framework that Uses Computers Like a Human
- MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses
- LLM With Tools: A Survey
- Interpreting Multi-band Galaxy Observations with Large Language Model-Based Agents
- HoneyComb: A Flexible LLM-Based Agent System for Materials Science
- Geometry Informed Tokenization of Molecules for Language Model Generation
- Fragment and Geometry Aware Tokenization of Molecules for Structure-Based Drug Design Using Language Models
- Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-Based Decoding
- Inverse designing metamaterials with programmable nonlinear functional responses in graph space
- From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future
- MetaOpenFOAM: an LLM-based multi-agent framework for CFD
- An Autonomous GIS Agent Framework for Geospatial Data Retrieval
- CellAgent: An LLM-driven Multi-Agent Framework for Automated Single-cell Data Analysis
- AtomAgents: Alloy design and discovery through physics-aware multi-modal multi-agent artificial intelligence
- Efficient Evolutionary Search Over Chemical Space with Large Language Models
- A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery
- Domain-specific ReAct for physics-integrated iterative modeling: A case study of LLM agents for gas path analysis of gas turbines
- BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments
- LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery
- A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models
- Accurate structure prediction of biomolecular interactions with AlphaFold 3
- CRISPR-GPT for Agentic Automation of Gene-editing Experiments
- Large Language Model Agent as a Mechanical Designer
- ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models
- Empowering Biomedical Discovery with AI Agents
- LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models
- Large Language Model-Based Evolutionary Optimizer: Reasoning with elitism
- Data Interpreter: An LLM Agent For Data Science
- AgentMD: Empowering Language Agents for Risk Prediction with Large-Scale Clinical Tool Learning
- Toward a Team of AI-made Scientists for Scientific Discovery from Gene Expression Data
- ChemReasoner: Heuristic Search over a Large Language Model's Knowledge Space using Quantum-Chemical Feedback
- Understanding the planning of LLM agents: A survey
- Executable Code Actions Elicit Better LLM Agents
- ProtAgents: Protein discovery via large language model multi-agent collaborations combining physics and machine learning
- Developing ChemDFM as a large language foundation model for chemistry
- Large Language Model based Multi-Agents: A Survey of Progress and Challenges
- TrustLLM: Trustworthiness in Large Language Models
- ChartAssisstant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning
- Targeted materials discovery using Bayesian algorithm execution
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Autonomous chemical research with large language models
- An autonomous laboratory for the accelerated synthesis of inorganic materials
- ChartLlama: A Multimodal LLM for Chart Understanding and Generation
- MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning
- Chemist-X: Large Language Model-empowered Agent for Reaction Condition Recommendation in Chemical Synthesis
- Large Language Models are Zero Shot Hypothesis Proposers
- Large Language Models as Evolutionary Optimizers
- Large Language Models Cannot Self-Correct Reasoning Yet
- AIKernel Semantic DSL Compiler and Deterministic Agent Execution Architecture
- Scientific discovery in the age of artificial intelligence
- Artificial Intelligence for Science in Quantum, Atomistic, and Continuum Systems
- Large Language Models
- De novo design of protein structure and function with RFdiffusion
- ClinicalGPT: Large Language Models Finetuned with Diverse Medical Data and Comprehensive Evaluation
- QH9: A Quantum Hamiltonian Prediction Benchmark for QM9 Molecules
- Efficient and Equivariant Graph Networks for Predicting Quantum Hamiltonian
- SciMON: Scientific Inspiration Machines Optimized for Novelty
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
- Teaching Large Language Models to Self-Debug
- ChemCrow: Augmenting large-language models with chemistry tools
- Self-Refine: Iterative Refinement with Self-Feedback
- Augmented Language Models: a Survey
- Evolutionary-scale prediction of atomic-level protein structure with a language model
- ReAct: Synergizing Reasoning and Acting in Language Models
- Atlas: Few-shot Learning with Retrieval Augmented Language Models
- Robust deep learning–based protein sequence design using ProteinMPNN
- LAMMPS - a flexible simulation tool for particle-based materials modeling at the atomic, meso, and continuum scales
- Highly accurate protein structure prediction with AlphaFold
- Effective gene expression prediction from sequence by integrating long-range interactions
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Don't Stop Pretraining: Adapt Language Models to Domains and Tasks
- REALM: Retrieval-Augmented Language Model Pre-Training
- A Survey on Knowledge Graphs: Representation, Acquisition, and Applications
- The atomic simulation environment—a Python library for working with atoms
- Electric Field Effect in Atomically Thin Carbon Films
- Fast Parallel Algorithms for Short-Range Molecular Dynamics
- Advancing AI-Scientist Understanding: Multi-Agent LLMs with Interpretable Physics Reasoning
- Reinforced Generation of Combinatorial Structures: Hardness of Approximation
- ChemMiner: A Large Language Model Agent System for Chemical Literature Data Mining
Cited by
Related