2020/12/15 by Oleksandr Gurbych, Maksym Druchok, Gurbych, Oleksandr +5
Biochemistry, Genetics and Molecular Biology · Computer Science · Materials Science · #Chemical Physics (physics.chem-ph) #Computational Drug Discovery Methods #FOS: Computer and information sciences #FOS: Physical sciences #Machine Learning (cs.LG) #Machine Learning in Materials Science #Protein Structure and Dynamics
paper · pdf · doi:10.48550/arxiv.2012.08275
openalex publication_date 2020/12/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This study assesses the efficiency of several popular machine learning approaches in the prediction of molecular binding affinity: CatBoost, Graph Attention Neural Network, and Bidirectional Encoder Representations from Transformers. The models were trained to predict binding affinities in terms of inhibition constants Ki for pairs of proteins and small organic molecules. First two approaches use thoroughly selected physico-chemical features, while the third one is based on textual molecular representations - it is one of the first attempts to apply Transformer-based predictors for the binding affinity. We also discuss the visualization of attention layers within the Transformer approach in order to highlight the molecular sites responsible for interactions. All approaches are free from atomic spatial coordinates thus avoiding bias from known structures and being able to generalize for compounds with unknown conformations. The achieved accuracy for all suggested approaches prove their potential in high throughput screening.