2018/04/30 by Asan Agibetov, Agibetov, Asan, Matthias Samwald +1 · 1 citation
Biochemistry, Genetics and Molecular Biology · #Artificial Intelligence (cs.AI) #Bioinformatics and Genomic Networks #Biomedical Text Mining and Ontologies #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning in Bioinformatics
paper · pdf · doi:10.48550/arxiv.1804.11105
openalex publication_date 2018/04/30 · openalex created_date 2022/10/06 · openalex updated_date 2026/07/28
In this work we address the problem of fast and scalable learning of\nneuro-symbolic representations for general biological knowledge. Based on a\nrecently published comprehensive biological knowledge graph (Alshahrani, 2017)\nthat was used for demonstrating neuro-symbolic representation learning, we show\nhow to train fast (under 1 minute) log-linear neural embeddings of the\nentities. We utilize these representations as inputs for machine learning\nclassifiers to enable important tasks such as biological link prediction.\nClassifiers are trained by concatenating learned entity embeddings to represent\nentity relations, and training classifiers on the concatenated embeddings to\ndiscern true relations from automatically generated negative examples. Our\nsimple embedding methodology greatly improves on classification error compared\nto previously published state-of-the-art results, yielding a maximum increase\nof +0.28 F-measure and +0.22 ROC AUC scores for the most difficult\nbiological link prediction problem. Finally, our embedding approach is orders\nof magnitude faster to train (\≤ 1 minute vs. hours), much more economical\nin terms of embedding dimensions (d=50 vs. d=512), and naturally encodes\nthe directionality of the asymmetric biological relations, that can be\ncontrolled by the order with which we concatenate the embeddings.\n