2019/12/27 by Siddhartha Nuthakki, Sunil Neela, Nuthakki, Siddhartha +5
Biochemistry, Genetics and Molecular Biology · Computer Science · Health Professions · #Biomedical Text Mining and Ontologies #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Healthcare #Medical Coding and Health Information
paper · pdf · doi:10.48550/arxiv.1912.12397
openalex publication_date 2019/12/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Coding diagnosis and procedures in medical records is a crucial process in\nthe healthcare industry, which includes the creation of accurate billings,\nreceiving reimbursements from payers, and creating standardized patient care\nrecords. In the United States, Billing and Insurance related activities cost\naround 471 billion in 2012 which constitutes about 25% of all the U.S hospital\nspending. In this paper, we report the performance of a natural language\nprocessing model that can map clinical notes to medical codes, and predict\nfinal diagnosis from unstructured entries of history of present illness,\nsymptoms at the time of admission, etc. Previous studies have demonstrated that\ndeep learning models perform better at such mapping when compared to\nconventional machine learning models. Therefore, we employed state-of-the-art\ndeep learning method, ULMFiT on the largest emergency department clinical notes\ndataset MIMIC III which has 1.2M clinical notes to select for the top-10 and\ntop-50 diagnosis and procedure codes. Our models were able to predict the\ntop-10 diagnoses and procedures with 80.3% and 80.5% accuracy, whereas the\ntop-50 ICD-9 codes of diagnosis and procedures are predicted with 70.7% and\n63.9% accuracy. Prediction of diagnosis and procedures from unstructured\nclinical notes benefit human coders to save time, eliminate errors and minimize\ncosts. With promising scores from our present model, the next step would be to\ndeploy this on a small-scale real-world scenario and compare it with human\ncoders as the gold standard. We believe that further research of this approach\ncan create highly accurate predictions that can ease the workflow in a clinical\nsetting.\n