ChemCrow: Augmenting large-language models with chemistry tools
2023/04/11 by Andres M Bran, Sam Cox, Bran, Andres M +9 · 2 voices · 121 citations
Materials Science · #Machine Learning in Materials Science #physics.chem-ph #stat.ML
paper · pdf · doi:10.48550/arxiv.2304.05376
openalex publication_date 2023/04/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Over the last decades, excellent computational chemistry tools have been developed. Integrating them into a single platform with enhanced accessibility could help reaching their full potential by overcoming steep learning curves. Recently, large-language models (LLMs) have shown strong performance in tasks across domains, but struggle with chemistry-related problems. Moreover, these models lack access to external knowledge sources, limiting their usefulness in scientific applications. In this study, we introduce ChemCrow, an LLM chemistry agent designed to accomplish tasks across organic synthesis, drug discovery, and materials design. By integrating 18 expert-designed tools, ChemCrow augments the LLM performance in chemistry, and new capabilities emerge. Our agent autonomously planned and executed the syntheses of an insect repellent, three organocatalysts, and guided the discovery of a novel chromophore. Our evaluation, including both LLM and expert assessments, demonstrates ChemCrow's effectiveness in automating a diverse set of chemical tasks. Surprisingly, we find that GPT-4 as an evaluator cannot distinguish between clearly wrong GPT-4 completions and Chemcrow's performance. Our work not only aids expert chemists and lowers barriers for non-experts, but also fosters scientific advancement by bridging the gap between experimental and computational chemistry.
Cited by
- MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing
- Model Gateway: Management Platform for Model-Driven Drug Discovery
- Agents in the Wild: Where Research Meets Deployment
- Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis
- Counting Cycles with AI: Counting Cycles with AI: Computationally Efficient Equivalent Forms with Applications
- Schema-Bound LLM Control of Scientific Instrumentation through Model Context Protocol Skills
- Grounded verification of chemical and materials reasoning: detection is the bottleneck
- HEPTAPOD: Orchestrating High Energy Physics Workflows Towards Autonomous Agency
- BrainPilot: Automating Brain Discovery with Agentic Research
- SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery
- AIMS: An uncertainty-aware AI experimentalist for quantum matter
- CADAQUES: A Cost-Aware Dual Architecture for Query-Efficient Autonomous Discovery
- A hierarchical memory architecture overcomes context limits in long-horizon multi-agent computational modeling
- Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science
- Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers
- CLINB: A Climate Intelligence Benchmark for Foundational Models
- AISSISTANT: Human-AI Collaborative Review and Perspective Research Workflows in Data Science
- 34 Examples of LLM Applications in Materials Science and Chemistry: Towards Automation, Assistants, Agents, and Accelerated Scientific Discovery
- El Agente: An autonomous agent for quantum chemistry
- When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
- How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?
- Materealize: a multi-agent deliberation system for end-to-end material design and synthesis
- SCTA: An Agentic Framework for Stable and Interpretable Target Gene Discovery from Single-Cell RNA Sequencing
- Performance of AI agents based on reasoning language models on ALD process optimization tasks
- Lexical discovery in unknown environments orchestrated by Large Language Models
- A Plan Reuse Mechanism for LLM-Driven Agent
- ChemATP: A Training-Free Chemical Reasoning Framework for Large Language Models
- ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
- Scalable Agentic Reasoning for Designing Biologics Targeting Intrinsically Disordered Proteins
- Dual-Axis RCCL: Representation-Complete Convergent Learning for Organic Chemical Space
- A Scientific Reasoning Model for Organic Synthesis Procedure Generation
- MedAI: Evaluating TxAgent's Therapeutic Agentic Reasoning in the NeurIPS CURE-Bench Competition
- From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models
- DynaMate: An Autonomous Agent for Protein-Ligand Molecular Dynamics Simulations
- WisPaper: Your AI Scholar Search Engine
- auto-psych: Automating the science of mind using agent-driven theory discovery and experimentation
- Process-Centric Analysis of Agentic Software Systems
- SynthStrategy: Extracting and Formalizing Latent Strategic Insights from LLMs in Organic Chemistry
- Hierarchical AI-Meteorologist: LLM-Agent System for Multi-Scale and Explainable Weather Forecast Reporting
- Solving Context Window Overflow in AI Agents
- CLIMATEAGENT: Multi-Agent Orchestration for Complex Climate Data Science Workflows
- Cross-Disciplinary Knowledge Retrieval and Synthesis: A Compound AI Architecture for Scientific Discovery
- Developing an AI Course for Synthetic Chemistry Students
- PrefixGPT: Prefix Adder Optimization by a Generative Pre-trained Transformer
- ChemVTS-Bench: Evaluating Visual-Textual-Symbolic Reasoning of Multimodal Large Language Models in Chemistry
- Teaching According to Students' Aptitude: Personalized Mathematics Tutoring via Persona-, Memory-, and Forgetting-Aware LLMs
- ChEmREF: Evaluating Language Model Readiness for Chemical Emergency Response
- Towards autonomous quantum physics research using LLM agents with access to intelligent tools
- Knowledge-Augmented Long-CoT Generation for Complex Biomolecular Reasoning
- AgenticSciML: Collaborative Multi-Agent Systems for Emergent Discovery in Scientific Machine Learning
- TinyChemVL: Advancing Chemical Vision-Language Models via Efficient Visual Token Reduction and Complex Reaction Tasks
- A Hierarchical Multi-Agent System for Autonomous Discovery in Geoscientific Data Archives
- Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments
- NISPO: Open-source IUPAC name generation tool
- CatPath‐GPT: A Mixture of Experts System for Computational Catalyst Design
- Atom-anchored LLMs speak Chemistry: A Retrosynthesis Demonstration
- Small Language Models Offer Significant Potential for Science Community
- ComProScanner: A multi-agent based framework for composition-property structured data extraction from scientific literature
- Foundation Models for Scientific Discovery: From Paradigm Enhancement to Paradigm Transition
- Multi-Agent Design Assistant for the Simulation of Inertial Fusion Energy
- FML-bench: A Benchmark for Automatic ML Research Agents Highlighting the Importance of Exploration Breadth
- Reasoning-Enhanced Large Language Models for Molecular Property Prediction
- MetaboT: An LLM-based Multi-Agent Frameworkfor Interactive Analysis of Mass SpectrometryMetabolomics Knowledge Graphs
- The Last Human-Written Paper: Agent-Native Research Artifacts
- Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics
- Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs
- oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning
- AgentAsk: Multi-Agent Systems Need to Ask
- Hypothesis Hunting with Evolving Networks of Autonomous Scientific Agents
- Auto-Stega: An Agent-Driven System for Lifelong Strategy Evolution in LLM-Based Text Steganography
- Expanding the Action Space of LLMs to Reason Beyond Language
- RareAgent: Self-Evolving Reasoning for Drug Repurposing in Rare Diseases
- Zephyrus: An Agentic Framework for Weather Science
- Allocation of Parameters in Transformers
- JoyAgent-JDGenie: Technical Report on the GAIA
- Speak to a Protein: An Interactive Multimodal Co-Scientist for Protein Analysis
- LLM Agents for Knowledge Discovery in Atomic Layer Processing
- Agentic Exploration of Physics Models
- Agentic Services Computing
- Mechanisms of Matter: Language Inferential Benchmark on Physicochemical Hypothesis in Materials Synthesis
- TusoAI: Agentic Optimization for Scientific Methods
- From What to Why: A Multi-Agent System for Evidence-based Chemical Reaction Condition Reasoning
- AOT*: Efficient Synthesis Planning via LLM-Empowered AND-OR Tree Search
- A Genetic Algorithm for Navigating Synthesizable Molecular Spaces
- LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?
- Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design Findings
- LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents
- Can AI Follow In Einstein's Footsteps?
- OSCAgent: Accelerating the Discovery of Organic Solar Cells with LLM Agents
- Spacer: Towards Engineered Scientific Inspiration
- Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
- Agentic AI for Multi-Stage Physics Experiments at a Large-Scale User Facility Particle Accelerator
- ChemOrch: Empowering LLMs with Chemical Intelligence via Synthetic Instructions
- OnlineMate: An LLM-Based Multi-Agent Companion System for Cognitive Support in Online Learning
- Foam-Agent 2.0: An End-to-End Composable Multi-Agent Framework for Automating CFD Simulation in OpenFOAM
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- MatSKRAFT: A framework for large-scale materials knowledge extraction from scientific tables
- Towards Fully Automated Molecular Simulations: Multi-Agent Framework for Simulation Setup and Force Field Extraction
- The (R)evolution of Scientific Workflows in the Agentic AI Era: Towards Autonomous Science
- Agents of Discovery
- Language Native Lightly Structured Databases for Large Language Model Driven Composite Materials Research
- The Need for Verification in AI-Driven Scientific Discovery
- Operating advanced scientific instruments with AI agents that learn on the job
- Osprey: Production-Ready Agentic AI for Safety-Critical Control Systems
- LARC: Towards Human-level Constrained Retrosynthesis Planning through an Agentic Framework
- The Rise of Generative AI for Metal-Organic Framework Design and Synthesis
- FROGENT: An End-to-End Full-process Drug Design Agent
- Retro-Expert: Collaborative Reasoning for Interpretable Retrosynthesis
- What are the limits to biomedical research acceleration through general-purpose AI?
- Multi-agent systems for chemical engineering: A review and perspective
- Towards Experience-Centered AI: A Framework for Integrating Lived Experience in Design and Development
- Simulating Human-Like Learning Dynamics with LLM-Empowered Agents
- Large Language Models Transform Organic Synthesis From Reaction Prediction to Automation
- LLM-based Multi-Agent Copilot for Quantum Sensor
- An Auditable Agent Platform For Automated Molecular Optimisation
- Autonomous Inorganic Materials Discovery via Multi-Agent Physics-Aware Scientific Reasoning
- Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications
- Graph-Augmented Large Language Model Agents: Current Progress and Future Prospects
- SciToolAgent: A Knowledge Graph-Driven Scientific Agent for Multi-Tool Integration
- Enhancing Molecular Structure Elucidation with Reasoning-Capable LLMs
- Innovator: Scientific Continued Pretraining with Fine-grained MoE Upcycling
Discussions
Related