2019/11/07 by Max Westphal, Antonia Zapf, Westphal, Max +3
Mathematics · Computer Science · #Statistical Methods in Clinical Trials #Machine Learning in Healthcare #Statistical Methods and Inference
paper · pdf · doi:10.48550/arxiv.1911.02982
Major advances have been made regarding the utilization of artificial\nintelligence in health care. In particular, deep learning approaches have been\nsuccessfully applied for automated and assisted disease diagnosis and prognosis\nbased on complex and high-dimensional data. However, despite all justified\nenthusiasm, overoptimistic assessments of predictive performance are still\ncommon. Automated medical testing devices based on machine-learned prediction\nmodels should thus undergo a throughout evaluation before being implemented\ninto clinical practice. In this work, we propose a multiple testing framework\nfor (comparative) phase III diagnostic accuracy studies with sensitivity and\nspecificity as co-primary endpoints. Our approach challenges the frequent\nrecommendation to strictly separate model selection and evaluation, i.e. to\nonly assess a single diagnostic model in the evaluation study. We show that our\nparametric simultaneous test procedure asymptotically allows strong control of\nthe family-wise error rate. Moreover, we demonstrate in extensive simulation\nstudies that our multiple testing strategy on average leads to a better final\ndiagnostic model and increased statistical power. To plan such studies, we\npropose a Bayesian approach to determine the optimal number of models to\nevaluate. For this purpose, our algorithm optimizes the expected final model\nperformance given previous (hold-out) data from the model development phase. We\nconclude that an assessment of multiple promising diagnostic models in the same\nevaluation study has several advantages when suitable adjustments for multiple\ncomparisons are implemented.\n