2017/11/14 by Madhumita Sushil, Simon Šuster, Sushil, Madhumita +5
Biochemistry, Genetics and Molecular Biology · Computer Science · Medicine · #Biomedical Text Mining and Ontologies #Colorectal Cancer Screening and Detection #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning in Healthcare
paper · pdf · doi:10.48550/arxiv.1711.05198
openalex publication_date 2017/11/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We have two main contributions in this work: 1. We explore the usage of a stacked denoising autoencoder, and a paragraph vector model to learn task-independent dense patient representations directly from clinical notes. We evaluate these representations by using them as features in multiple supervised setups, and compare their performance with those of sparse representations. 2. To understand and interpret the representations, we explore the best encoded features within the patient representations obtained from the autoencoder model. Further, we calculate the significance of the input features of the trained classifiers when we use these pretrained representations as input.