2024/01/12 by Lingchao Mao, Mao, Lingchao, Hairong Wang +11
Biochemistry, Genetics and Molecular Biology · Medicine · #92B99 #Artificial Intelligence (cs.AI) #Bioinformatics and Genomic Networks #FOS: Computer and information sciences #Gene expression and cancer classification #Machine Learning (cs.LG) #Radiomics and Machine Learning in Medical Imaging
paper · pdf · doi:10.48550/arxiv.2401.06406
openalex publication_date 2024/01/12 · openalex created_date 2024/01/16 · openalex updated_date 2026/07/28
Cancer remains one of the most challenging diseases to treat in the medical field. Machine learning has enabled in-depth analysis of rich multi-omics profiles and medical imaging for cancer diagnosis and prognosis. Despite these advancements, machine learning models face challenges stemming from limited labeled sample sizes, the intricate interplay of high-dimensionality data types, the inherent heterogeneity observed among patients and within tumors, and concerns about interpretability and consistency with existing biomedical knowledge. One approach to surmount these challenges is to integrate biomedical knowledge into data-driven models, which has proven potential to improve the accuracy, robustness, and interpretability of model results. Here, we review the state-of-the-art machine learning studies that adopted the fusion of biomedical knowledge and data, termed knowledge-informed machine learning, for cancer diagnosis and prognosis. Emphasizing the properties inherent in four primary data types including clinical, imaging, molecular, and treatment data, we highlight modeling considerations relevant to these contexts. We provide an overview of diverse forms of knowledge representation and current strategies of knowledge integration into machine learning pipelines with concrete examples. We conclude the review article by discussing future directions to advance cancer research through knowledge-informed machine learning.