2018/07/06 by M. Tarık Altuncu, Erik Mayer, Altuncu, M. Tarik +5
Biochemistry, Genetics and Molecular Biology · Computer Science · Decision Sciences · #Biomedical Text Mining and Ontologies #Computation and Language (cs.CL) #Data Quality and Management #FOS: Computer and information sciences #FOS: Mathematics #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Machine Learning in Healthcare #Social and Information Networks (cs.SI) #Spectral Theory (math.SP) #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1807.02599
openalex publication_date 2018/07/06 · openalex created_date 2022/08/04 · openalex updated_date 2026/07/28
Electronic Healthcare Records contain large volumes of unstructured data,\nincluding extensive free text. Yet this source of detailed information often\nremains under-used because of a lack of methodologies to extract interpretable\ncontent in a timely manner. Here we apply network-theoretical tools to analyse\nfree text in Hospital Patient Incident reports from the National Health\nService, to find clusters of documents with similar content in an unsupervised\nmanner at different levels of resolution. We combine deep neural network\nparagraph vector text-embedding with multiscale Markov Stability community\ndetection applied to a sparsified similarity graph of document vectors, and\nshowcase the approach on incident reports from Imperial College Healthcare NHS\nTrust, London. The multiscale community structure reveals different levels of\nmeaning in the topics of the dataset, as shown by descriptive terms extracted\nfrom the clusters of records. We also compare a posteriori against hand-coded\ncategories assigned by healthcare personnel, and show that our approach\noutperforms LDA-based models. Our content clusters exhibit good correspondence\nwith two levels of hand-coded categories, yet they also provide further medical\ndetail in certain areas and reveal complementary descriptors of incidents\nbeyond the external classification taxonomy.\n