vix.ing · top · new · best · stats

Using Ontologies To Improve Performance In Massively Multi-label Prediction Models

2019/05/28 by Ethan Steinberg, Peter J. Liu, Steinberg, Ethan +1 · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · Mathematics · #Artificial Intelligence (cs.AI) #Artificial intelligence #Artificial neural network #Biology #Biomedical Text Mining and Ontologies #Computer science #Data mining #FOS: Computer and information sciences #Function (biology) #Gene #Gene ontology #Layer (electronics) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning in Bioinformatics #Machine learning #Massively parallel #Multi-label classification #Ontology #Parallel computing #Topic Modeling #cs.AI #cs.LG #stat.ML

paper · pdf · doi:10.48550/arxiv.1905.12126

published in arXiv (Cornell University) (Cornell University)

arxiv created 2019/05/28 · openalex publication_date 2019/05/28 · arxiv updated 2019/05/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Massively multi-label prediction/classification problems arise in environments like health-care or biology where very precise predictions are useful. One challenge with massively multi-label problems is that there is often a long-tailed frequency distribution for the labels, which results in few positive examples for the rare labels. We propose a solution to this problem by modifying the output layer of a neural network to create a Bayesian network of sigmoids which takes advantage of ontology relationships between the labels to help share information between the rare and the more common labels. We apply this method to the two massively multi-label tasks of disease prediction (ICD-9 codes) and protein function prediction (Gene Ontology terms) and obtain significant improvements in per-label AUROC and average precision for less common labels.

Citations

Related