vix.ing · top · new · best · stats · spec

Validity problems in clinical machine learning by indirect data labeling using consensus definitions

2023/11/06 by Michael Hagmann, Hagmann, Michael, Shigehiko Schamoni +3 · 1 citation
Computer Science · Medicine · #Applications (stat.AP) #Explainable Artificial Intelligence (XAI) #FOS: Biological sciences #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning in Healthcare #Quantitative Methods (q-bio.QM) #Sepsis Diagnosis and Treatment

paper · pdf · doi:10.48550/arxiv.2311.03037

openalex publication_date 2023/11/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We demonstrate a validity problem of machine learning in the vital application area of disease diagnosis in medicine. It arises when target labels in training data are determined by an indirect measurement, and the fundamental measurements needed to determine this indirect measurement are included in the input data representation. Machine learning models trained on this data will learn nothing else but to exactly reconstruct the known target definition. Such models show perfect performance on similarly constructed test data but will fail catastrophically on real-world examples where the defining fundamental measurements are not or only incompletely available. We present a general procedure allowing identification of problematic datasets and black-box machine learning models trained on them, and exemplify our detection procedure on the task of early prediction of sepsis.

Cited by

Related