2023/06/13 by Yin Fang, Fang, Yin, Xiaozhuan Liang +13 · 14 citations
Biochemistry, Genetics and Molecular Biology · Materials Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computational Engineering #FOS: Biological sciences #FOS: Computer and information sciences #Finance #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Machine Learning in Bioinformatics #Machine Learning in Materials Science #Protein Structure and Dynamics #Quantitative Methods (q-bio.QM) #and Science (cs.CE)
paper · pdf · doi:10.48550/arxiv.2306.08018
openalex publication_date 2023/06/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Large Language Models (LLMs), with their remarkable task-handling capabilities and innovative outputs, have catalyzed significant advancements across a spectrum of fields. However, their proficiency within specialized domains such as biomolecular studies remains limited. To address this challenge, we introduce Mol-Instructions, a comprehensive instruction dataset designed for the biomolecular domain. Mol-Instructions encompasses three key components: molecule-oriented instructions, protein-oriented instructions, and biomolecular text instructions. Each component aims to improve the understanding and prediction capabilities of LLMs concerning biomolecular features and behaviors. Through extensive instruction tuning experiments on LLMs, we demonstrate the effectiveness of Mol-Instructions in enhancing large models' performance in the intricate realm of biomolecular studies, thus fostering progress in the biomolecular research community. Mol-Instructions is publicly available for ongoing research and will undergo regular updates to enhance its applicability.