2017/11/28 by Hai Nguyen, Nguyen, Hai, Shin‐ichi Maeda +3
Biochemistry, Genetics and Molecular Biology · Computer Science · Materials Science · #Computational Drug Discovery Methods #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning in Bioinformatics #Machine Learning in Materials Science
paper · pdf · doi:10.48550/arxiv.1711.10168
openalex publication_date 2017/11/28 · openalex created_date 2017/12/04 · openalex updated_date 2026/07/28
With the rapid increase of compound databases available in medicinal and material science, there is a growing need for learning representations of molecules in a semi-supervised manner. In this paper, we propose an unsupervised hierarchical feature extraction algorithm for molecules (or more generally, graph-structured objects with fixed number of types of nodes and edges), which is applicable to both unsupervised and semi-supervised tasks. Our method extends recently proposed Paragraph Vector algorithm and incorporates neural message passing to obtain hierarchical representations of subgraphs. We applied our method to an unsupervised task and demonstrated that it outperforms existing proposed methods in several benchmark datasets. We also experimentally showed that semi-supervised tasks enhanced predictive performance compared with supervised ones with labeled molecules only.