2026/06/04 by Alexander David Gibson, Nicole M White, Gary S Collins +1 · 2 voices
Medicine · Decision Sciences · #Artificial Intelligence in Healthcare and Education #Scientific Computing and Data Management #Meta-analysis and systematic reviews
paper · doi:10.1186/s12916-026-04981-y
BACKGROUND: Clinical prediction models are often created using large routinely collected datasets. It is essential that prediction models are developed with appropriate data and methods and transparently reported to ensure that decisions are based on reliable predictions. Kaggle is a popular competition and data repository website where users learn and apply analysis skills on a range of datasets. METHODS: We identified two large, publicly available Kaggle datasets, on stroke and diabetes, that lack clear data provenance, but are widely used in clinical prediction models in peer reviewed publications. We used exploratory analyses to examine the quality of data and reporting of information using nine items from the TRIPOD+AI statement checklist. RESULTS: Data provenance assessment using nine TRIPOD+AI items revealed major deficiencies, with minimal details for either dataset including no information on when, where, why or how the data were collected. The authenticity of both datasets could not be verified and have no reliable provenance of authenticity and should not be used for informing research or practice. From these two datasets, we found 125 clinical prediction model studies. Three prediction models had evidence of use in clinical practice, one model was cited in a medical device patent, and the models were cited in 86 review articles. CONCLUSIONS: We recommend that journals and data repositories mandate data provenance reporting to safeguard published research. Prediction models based solely on inauthentic or unreliable datasets should never be used to directly inform decisions on patient care.