2019/09/26 by Luke Oakden‐Rayner, Jared Dunnmon, Oakden-Rayner, Luke +5 · 18 citations
Medicine · Computer Science · Mathematics · #Radiomics and Machine Learning in Medical Imaging #AI in cancer detection #Statistical Methods and Inference
paper · pdf · doi:10.48550/arxiv.1909.12475
Machine learning models for medical image analysis often suffer from poor\nperformance on important subsets of a population that are not identified during\ntraining or testing. For example, overall performance of a cancer detection\nmodel may be high, but the model still consistently misses a rare but\naggressive cancer subtype. We refer to this problem as hidden stratification,\nand observe that it results from incompletely describing the meaningful\nvariation in a dataset. While hidden stratification can substantially reduce\nthe clinical efficacy of machine learning models, its effects remain difficult\nto measure. In this work, we assess the utility of several possible techniques\nfor measuring and describing hidden stratification effects, and characterize\nthese effects on multiple medical imaging datasets. We find evidence that\nhidden stratification can occur in unidentified imaging subsets with low\nprevalence, low label quality, subtle distinguishing features, or spurious\ncorrelates, and that it can result in relative performance differences of over\n20% on clinically important subsets. Finally, we explore the clinical\nimplications of our findings, and suggest that evaluation of hidden\nstratification should be a critical component of any machine learning\ndeployment in medical imaging.\n