Knowledge-Augmented Long-CoT Generation for Complex Biomolecular Reasoning
2025/11/11 by Xiang Zhuang, Lyu, Tianwen, Zhuang, Xiang +11 · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning in Bioinformatics #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2511.08024
openalex publication_date 2025/11/11 · openalex created_date 2025/11/13 · openalex updated_date 2026/07/28
Abstract
Understanding complex biomolecular mechanisms requires multi-step reasoning across molecular interactions, signaling cascades, and metabolic pathways. While large language models(LLMs) show promise in such tasks, their application to biomolecular problems is hindered by logical inconsistencies and the lack of grounding in domain knowledge. Existing approaches often exacerbate these issues: reasoning steps may deviate from biological facts or fail to capture long mechanistic dependencies. To address these challenges, we propose a Knowledge-Augmented Long-CoT Reasoning framework that integrates LLMs with knowledge graph-based multi-hop reasoning chains. The framework constructs mechanistic chains via guided multi-hop traversal and pruning on the knowledge graph; these chains are then incorporated into supervised fine-tuning to improve factual grounding and further refined with reinforcement learning to enhance reasoning reliability and consistency. Furthermore, to overcome the shortcomings of existing benchmarks, which are often restricted in scale and scope and lack annotations for deep reasoning chains, we introduce PrimeKGQA, a comprehensive benchmark for biomolecular question answering. Experimental results on both PrimeKGQA and existing datasets demonstrate that although larger closed-source models still perform well on relatively simple tasks, our method demonstrates clear advantages as reasoning depth increases, achieving state-of-the-art performance on multi-hop tasks that demand traversal of structured biological knowledge. These findings highlight the effectiveness of combining structured knowledge with advanced reasoning strategies for reliable and interpretable biomolecular reasoning.
Citations
- KG-o1: Enhancing Multi-hop Question Answering in Large Language Models via Knowledge Graph Integration
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
- Training a Scientific Reasoning Model for Chemistry
- Cell-o1: Training LLMs to Solve Single-Cell Reasoning Puzzles with Reinforcement Learning
- BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model
- Qwen3 Technical Report
- MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs
- BioMaze: Benchmarking and Enhancing Large Language Models for Biological Pathway Reasoning
- Competitive Programming with Large Reasoning Models
- s1: Simple test-time scaling
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
- KG4Diagnosis: A Hierarchical Multi-Agent LLM Framework with Knowledge Graph Enhancement for Medical Diagnosis
- Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems
- Towards Next-Generation Medical Agent: How o1 is Reshaping Decision-Making in Medical Scenarios
- GPT-4o System Card
- Graph-constrained Reasoning: Faithful Reasoning on Knowledge Graphs with Large Language Models
- Adaptive Reasoning and Acting in Medical Language Agents
- Advancing biomolecular understanding and design following human instructions
- MKGL: Mastery of a Three-Word Language
- Large Language Models in Drug Discovery and Development: From Disease Mechanisms to Clinical Trials
- MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine
- CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis
- BiomedRAG: A Retrieval Augmented Large Language Model for Biomedicine
- STaRK: Benchmarking LLM Retrieval on Textual and Relational Knowledge Bases
- ProLLM: Protein Chain-of-Thoughts Enhanced LLM for Protein-Protein Interaction Prediction
- Progress and Opportunities of Foundation Models in Bioinformatics
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- Scientific Large Language Models: A Survey on Biological & Chemical Domains
- Biomedical knowledge graph-optimized prompt generation for large language models
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- BioBridge: Bridging Biomedical Foundation Models via Knowledge Graphs
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
- Comparative Performance Evaluation of Large Language Models for Extracting Molecular Interactions and Pathway Knowledge
- Self-Polish: Enhance Reasoning in Large Language Models via Problem Refinement
- ChemCrow: Augmenting large-language models with chemistry tools
- Retrieved Sequence Augmentation for Protein Representation Learning
- Specializing Smaller Language Models towards Multi-Step Reasoning
- Domain-Agnostic Molecular Generation with Chemical Feedback
- Large Language Models Are Reasoning Teachers
- Teaching Small Language Models to Reason
- Complexity-Based Prompting for Multi-Step Reasoning
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
- Large Language Models are Zero-Shot Reasoners
- Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Neural Multi-Hop Reasoning With Logical Rules on Biomedical Knowledge Graphs
- PRAG: Paninian Retrieval-Augmented Generation for Safety-Critical Medical Question Answering
- PubMedQA: A Dataset for Biomedical Research Question Answering
- OpenAI o1 System Card
Cited by
Related