2026/03/19 by Sijian Fan, Liyan Xiong, Dayuan Wang +2 · 1 voice
Biochemistry, Genetics and Molecular Biology · Computer Science · Materials Science · Mathematics · #Bayes' theorem #Bayesian network #Bayesian probability #Computational Drug Discovery Methods #Drug discovery #Feature selection #Latent variable #Machine Learning in Bioinformatics #Machine Learning in Materials Science #Relevance (law) #Selection (genetic algorithm) #Variable (mathematics) #cs.LG #stat.ME
paper · pdf · doi:10.48550/arxiv.2603.18957
openalex publication_date 2026/03/19 · arxiv published 2026/03/19 · arxiv updated 2026/03/19 · openalex created_date 2026/03/21 · openalex updated_date 2026/07/28
Recent advances in drug discovery have demonstrated that incorporating side information (e.g., chemical properties about drugs and genomic information about diseases) often greatly improves prediction performance. However, these side features can vary widely in relevance and are often noisy and high-dimensional. We propose Bayesian Variable Selection-Guided Inductive Matrix Completion (BVSIMC), a new Bayesian model that enables variable selection from side features in drug discovery. By learning sparse latent embeddings, BVSIMC improves both predictive accuracy and interpretability. We validate our method through simulation studies and two drug discovery applications: 1) prediction of drug resistance in Mycobacterium tuberculosis, and 2) prediction of new drug-disease associations in computational drug repositioning. On both synthetic and real data, BVSIMC outperforms several other state-of-the-art methods in terms of prediction. In our two real examples, BVSIMC further reveals the most clinically meaningful side features.