Deciphering Scientific Reasoning Steps from Outcome Data for Molecule Optimization
2026/03/13 by Zequn Liu, Kehan Wu, Shufang Xie +5 · 1 voice
Biochemistry, Genetics and Molecular Biology · Computer Science · #cs.AI #cs.LG #q-bio.BM
paper · pdf · doi:10.48550/arxiv.2603.20262
Abstract
Emerging reasoning models hold promise for automating scientific discovery. However, their training is hindered by a critical supervision gap: experimental outcomes are abundant, whereas intermediate reasoning steps are rarely documented at scale. To bridge this gap, we propose DESRO, a framework for deciphering scientific reasoning from outcomes. By analyzing shared patterns and key differences within grouped data, a large language model (LLM) can recover the underlying logic. We instantiate this framework in molecule optimization, a pivotal stage in drug discovery that traditionally relies on the iterative reasoning of medicinal chemists. Across 2.3 million molecular property records, our framework infers optimization rationales by grouping molecules with shared fragments, then using an LLM to analyze how structural variations correlate with property differences. Based on the derived data, we train a model that conducts molecule optimization through an interpretable reasoning process. DESRO achieves the highest success rates on 15 out of 18 tasks, spanning both single- and multi-property optimization of bioactivity and ADMET properties. The reasoning process enables robust generalization to out-of-distribution scenarios, including novel property combinations, unseen biological targets, and unseen properties defined solely by natural language descriptions. In retrospective case studies under strict temporal splits, the model autonomously reconstructs expert-level lead optimization trajectories. Additionally, our framework extends beyond molecule optimization to reaction ligand selection. Our results establish deciphering reasoning steps from outcome data as a viable paradigm for enabling scientific reasoning, providing a scalable approach to accelerate scientific discovery.
Citations
- Chem-R: Learning to Reason as a Chemist
- Coder as Editor: Code-driven Interpretable Molecular Optimization
- POLO: Preference-Guided Multi-Turn Reinforcement Learning for Lead Optimization
- SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines
- Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications
- ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge
- Conditional Chemical Language Models are Versatile Tools in Drug Discovery
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Training a Scientific Reasoning Model for Chemistry
- ChemAU: Harness the Reasoning of LLMs in Chemical Research with Adaptive Uncertainty Estimation
- BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model
- Large Language Models for Controllable Multi-property Multi-objective Molecule Optimization
- Beyond Chemical QA: Evaluating LLM's Chemical Reasoning with Modular Chemical Operations
- MT-Mol:Multi Agent System with Tool-based Reasoning for Molecular Optimization
- Qwen3 Technical Report
- PharmAgents: Building a Virtual Pharma with Large Language Model Agents
- Collaborative Expert LLMs Guided Multi-Objective Molecular Optimization
- InversionGNN: A Dual Path Network for Multi-Property Molecular Optimization
- DeepSeek vs. ChatGPT vs. Claude: A Comparative Study for Scientific Computing and Scientific Machine Learning Tasks
- GeLLMO: Generalizing Large Language Models for Multi-property Molecule Optimization
- Nature Language Model: Deciphering the Language of Nature for Scientific Discovery
- GenMol: A Drug Discovery Generalist with Discrete Diffusion
- Molecule Generation with Fragment Retrieval Augmentation
- Generative Artificial Intelligence for Navigating Synthesizable Chemical Space
- The Llama 3 Herd of Models
- Directly Optimizing for Synthesizability in Generative Molecular Design using Retrosynthesis Models
- LICO: Large Language Models for In-Context Molecular Optimization
- ACEGEN: Reinforcement learning of generative chemical agents for drug discovery
- LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Comprehensive, High-Quality Instruction Tuning Dataset
- Graph Diffusion Transformers for Multi-Conditional Molecular Generation
- DrugAssist: A Large Language Model for Molecule Optimization
- Autonomous chemical research with large language models
- ChatGPT-powered Conversational Drug Editing Using Retrieval and Domain Feedback
- Sample-efficient Multi-objective Molecular Optimization with GFlowNets
- De Novo Molecular Generation via Connection-aware Motif Mining
- Multi-modal Molecule Structure-text Model for Text-based Retrieval and Editing
- Reinforced Genetic Algorithm for Structure-based Drug Design
- PubChem 2023 update
- Exploring Chemical Space with Score-based Out-of-distribution Generation
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Local Latent Space Bayesian Optimization over Structured Inputs
- Amortized Tree Generation for Bottom-up Synthesis Planning and Synthesizable Molecular Design
- Hit and Lead Discovery with Explorative RL and Fragment-based Molecule Generation
- Differentiable Scaffolding Tree for Molecular Optimization
- MARS: Markov Molecular Sampling for Multi-objective Drug Discovery
- Therapeutics Data Commons: Machine Learning Datasets and Tasks for Drug Discovery and Development
- MIMOSA: Multi-constraint Molecule Sampling for Molecule Optimization
- Molecular Design in Synthetically Accessible Chemical Space via Deep Reinforcement Learning
- Augmenting Genetic Algorithms with Deep Neural Networks for Exploring the Chemical Space
- ChemBO: Bayesian Optimization of Small Organic Molecules with Synthesizable Recommendations
- GuacaMol: Benchmarking Models for De Novo Molecular Design
- Optimization of Molecules via Deep Reinforcement Learning
- Graph Convolutional Policy Network for Goal-Directed Molecular Graph Generation
- Population-based de novo molecule generation, using grammatical evolution
- Junction Tree Variational Autoencoder for Molecular Graph Generation
- Molecular De Novo Design through Deep Reinforcement Learning
- Computational Methods in Drug Discovery
- ChEMBL: a large-scale bioactivity database for drug discovery
- AutoDock Vina: Improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading
- OpenAI o1 System Card
Discussions
Related