2024/08/09 by Amish Mishra, Mishra, Amish, Francis C. Motta +1
Biochemistry, Genetics and Molecular Biology · Computer Science · Mathematics · #Algorithm #Artificial intelligence #Combinatorics #Computational Drug Discovery Methods #Computer science #Data Analysis #Data mining #FOS: Computer and information sciences #FOS: Physical sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine learning #Mathematics #Microbial Metabolic Engineering and Bioproduction #Pipeline (software) #Protein Structure and Dynamics #Stability (learning theory) #Statistics and Probability (physics.data-an) #Topological data analysis #Topology (electrical circuits)
paper · pdf · doi:10.48550/arxiv.2408.04847
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/08/09 · openalex created_date 2024/09/10 · openalex updated_date 2026/07/28
In this paper, we propose a data-driven method to learn interpretable topological features of biomolecular data and demonstrate the efficacy of parsimonious models trained on topological features in predicting the stability of synthetic mini proteins. We compare models that leverage automatically-learned structural features against models trained on a large set of biophysical features determined by subject-matter experts (SME). Our models, based only on topological features of the protein structures, achieved 92%-99% of the performance of SME-based models in terms of the average precision score. By interrogating model performance and feature importance metrics, we extract numerous insights that uncover high correlations between topological features and SME features. We further showcase how combining topological features and SME features can lead to improved model performance over either feature set used in isolation, suggesting that, in some settings, topological features may provide new discriminating information not captured in existing SME features that are useful for protein stability prediction.