2018/08/08 by Peter J. Liu, Liu, Peter J. · 1 citation
Computer Science · Health Professions · #Computation and Language (cs.CL) #Electronic Health Records Systems #FOS: Computer and information sciences #Machine Learning in Healthcare #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.1808.02622
preprint
arxiv created 2018/08/08 · openalex publication_date 2018/08/08 · arxiv updated 2018/08/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Clinicians spend a significant amount of time inputting free-form textual notes into Electronic Health Records (EHR) systems. Much of this documentation work is seen as a burden, reducing time spent with patients and contributing to clinician burnout. With the aspiration of AI-assisted note-writing, we propose a new language modeling task predicting the content of notes conditioned on past data from a patient's medical record, including patient demographics, labs, medications, and past notes. We train generative models using the public, de-identified MIMIC-III dataset and compare generated notes with those in the dataset on multiple measures. We find that much of the content can be predicted, and that many common templates found in notes can be learned. We discuss how such models can be useful in supporting assistive note-writing features such as error-detection and auto-complete.