2025/09/10 by Dong Han, Zizheng Ai, Han, Dong +34 · 1 citation
Computer Science · Materials Science · #Baseline (sea) #Bayesian optimization #Bayesian probability #Branch and bound #Computational Drug Discovery Methods #FOS: Computer and information sciences #Initialization #Linear subspace #Machine Learning (cs.LG) #Machine Learning and Data Classification #Machine Learning in Materials Science #Set (abstract data type) #Upper and lower bounds
paper · pdf · doi:10.48550/arxiv.2509.08736
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/09/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Bayesian optimization (BO) is a powerful tool for scientific discovery in chemistry, yet its efficiency is often hampered by the sparse experimental data and vast search space. Here, we introduce ChemBOMAS: a large language model (LLM)-enhanced multi-agent system that accelerates BO through synergistic data- and knowledge-driven strategies. Firstly, the data-driven strategy involves an 8B-scale LLM regressor fine-tuned on a mere 1% labeled samples for pseudo-data generation, robustly initializing the optimization process. Secondly, the knowledge-driven strategy employs a hybrid Retrieval-Augmented Generation approach to guide LLM in dividing the search space while mitigating LLM hallucinations. An Upper Confidence Bound algorithm then identifies high-potential subspaces within this established partition. Across the LLM-refined subspaces and supported by LLM-generated data, BO achieves the improvement of effectiveness and efficiency. Comprehensive evaluations across multiple scientific benchmarks demonstrate that ChemBOMAS set a new state-of-the-art, accelerating optimization efficiency by up to 5-fold compared to baseline methods.