vix.ing · top · new · best · stats · spec

Clinical Concept Embeddings Learned from Massive Sources of Multimodal\n Medical Data

2018/04/04 by Andrew L. Beam, Benjamin Kompa, Beam, Andrew L. +15 · 4 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Artificial Intelligence (cs.AI) #Biomedical Text Mining and Ontologies #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (stat.ML) #Natural Language Processing Techniques #Topic Modeling

paper · pdf · doi:10.48550/arxiv.1804.01486

openalex publication_date 2018/04/04 · openalex created_date 2022/10/04 · openalex updated_date 2026/07/28

Abstract

Word embeddings are a popular approach to unsupervised learning of word\nrelationships that are widely used in natural language processing. In this\narticle, we present a new set of embeddings for medical concepts learned using\nan extremely large collection of multimodal medical data. Leaning on recent\ntheoretical insights, we demonstrate how an insurance claims database of 60\nmillion members, a collection of 20 million clinical notes, and 1.7 million\nfull text biomedical journal articles can be combined to embed concepts into a\ncommon space, resulting in the largest ever set of embeddings for 108,477\nmedical concepts. To evaluate our approach, we present a new benchmark\nmethodology based on statistical power specifically designed to test embeddings\nof medical concepts. Our approach, called cui2vec, attains state-of-the-art\nperformance relative to previous methods in most instances. Finally, we provide\na downloadable set of pre-trained embeddings for other researchers to use, as\nwell as an online tool for interactive exploration of the cui2vec embeddings\n

Cited by

Related