Calibration: the Achilles heel of predictive analytics
2019/12/01 by On behalf of Topic Group ‘Evaluating diagnostic tests and prediction models’ of the STRATOS initiative, Ben Van Calster, David J. McLernon +3 · 1 voice · 36 citations
Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI) #Sepsis Diagnosis and Treatment
paper · pdf · doi:10.1186/s12916-019-1466-7
openalex publication_date 2019/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30
Abstract
BACKGROUND: The assessment of calibration performance of risk prediction models based on regression or more flexible machine learning algorithms receives little attention. MAIN TEXT: Herein, we argue that this needs to change immediately because poorly calibrated algorithms can be misleading and potentially harmful for clinical decision-making. We summarize how to avoid poor calibration at algorithm development and how to assess calibration at algorithm validation, emphasizing balance between model complexity and the available sample size. At external validation, calibration curves require sufficiently large samples. Algorithm updating should be considered for appropriate support of clinical practice. CONCLUSION: Efforts are required to avoid poor calibration when developing prediction models, to evaluate calibration when validating models, and to update models when indicated. The ultimate aim is to optimize the utility of predictive analytics for shared decision-making and patient counseling.
Citations
Cited by
- A comparison of hyperparameter tuning procedures for clinical prediction models: A simulation study
- Artificial intelligence for modelling infectious disease epidemics
- Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance
- Bellman Calibration for V-Learning in Offline Reinforcement Learning
- Stability of clinical prediction models developed using statistical or machine learning methods
- Harmonized Interpretable ECG Waveform Features for Robust Cross-Dataset Clinical Prediction
- WTNN: Weibull-Tailored Neural Networks for survival analysis
- A Critical Perspective on Finite Sample Conformal Prediction Theory in Medical Applications
- Non-parametric assessment of the calibration of individualized treatment effects
- Machine learning for violence prediction: a systematic review and critical appraisal
- Improving accuracy in the estimation of probable dementia in racially and ethnically diverse groups with penalized regression and transfer learning
- Validation of the Clinical Frailty Scale for predicting 90-day mortality in hospitalised older adults screened as at risk of nearing the end of life in Queensland, Australia: a multisite observational study
- The Mayo MGRS Prediction Tool calculates the risk of finding monoclonal gammopathy of renal significance in a kidney biopsy in patients with monoclonal gammopathy
- Navigating prevalence shifts in image analysis algorithm deployment
- Prognostic models predicting clinical outcomes in patients diagnosed with visceral leishmaniasis: a systematic review
- Cost-Efficient Early Diagnostic Tool for Lung Cancer: Explainable AI in Clinical Systems
- Predicting admission to and length of stay in intensive care units after general anesthesia: Time-dependent role of pre- and intraoperative data for clinical decision-making
- Progress, challenges, and pragmatic concessions in predicting relative risk of kidney survival in ARPKD
- Bayesian Sequential Modeling of Time-to-Urination for Dynamic ED Triage
- International multicenter validation of AI-driven ultrasound detection of ovarian cancer
- SoC-DT: Standard-of-Care Aligned Digital Twins for Patient-Specific Tumor Dynamics
- The harm of class imbalance corrections for risk prediction models: illustration and simulation using logistic regression
- Uncertainty of risk estimates from clinical prediction models: rationale, challenges, and approaches
- Limitations on Safe, Trusted, Artificial General Intelligence
- Artificial Intelligence Across the Surgical Oncology Continuum: Decision Support, Operative Intelligence, and a Translation-First Roadmap
- Antipsychotic-induced weight gain in psychosis: causal mediation analysis and feasibility study of causal actionable prediction model development using counterfactuals to target obesity
- Clinical prediction models and the multiverse of madness
- Artificial intelligence–based quantification of breast arterial calcifications to predict cardiovascular morbidity and mortality
- A predictor finding study found patient-reported outcomes improve the prediction of mortality of incident dialysis patients
- There is no such thing as a validated prediction model
- Prioritising deteriorating patients using time-to-event analysis: prediction model development and internal–external validation
- When the whole is greater than the sum of its parts: why machine learning and conventional statistics are complementary for predicting future health outcomes
- Predicting type 2 diabetes and testosterone effects in high-risk Australian men: development and external validation of a 2-year risk model
- PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods
- Prediction models for diagnosis and prognosis of covid-19: systematic review and critical appraisal
- Statistical Methods in Generative AI
Discussions
Related