2021/07/03 by Anna Fedyukova, Douglas E. V. Pires, Fedyukova, Anna +4
Computer Science · Medicine · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Healthcare #Sepsis Diagnosis and Treatment #Time Series Analysis and Forecasting #cs.LG
paper · pdf · doi:10.48550/arxiv.2107.10399
3 pages, 1 figure, Joint KDD 2021 Health Day and 2021 KDD Workshop on Applied Data Science for Healthcare, August 14-18, 2021
arxiv created 2021/07/03 · openalex publication_date 2021/07/03 · arxiv updated 2021/07/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The proliferation of early diagnostic technologies, including self-monitoring systems and wearables, coupled with the application of these technologies on large segments of healthy populations may significantly aggravate the problem of overdiagnosis. This can lead to unwanted consequences such as overloading health care systems and overtreatment, with potential harms to healthy individuals. The advent of machine-learning tools to assist diagnosis -- while promising rapid and more personalised patient management and screening -- might contribute to this issue. The identification of overdiagnosis is usually post hoc and demonstrated after long periods (from years to decades) and costly randomised control trials. In this paper, we present an innovative approach that allows us to preemptively detect potential cases of overdiagnosis during predictive model development. This approach is based on the combination of labels obtained from a prediction model and clustered medical trajectories, using sepsis in adults as a test case. This is one of the first attempts to quantify machine-learning induced overdiagnosis and we believe will serves as a platform for further development, leading to guidelines for safe deployment of computational diagnostic tools.