Calculating the sample size required for developing a clinical prediction model
2020/03/18 by Richard D Riley, Joie Ensor, Kym I E Snell +6 · 38 citations
Decision Sciences · Mathematics · Medicine · #Meta-analysis and systematic reviews #Sepsis Diagnosis and Treatment #Statistical Methods in Epidemiology
paper · pdf · doi:10.1136/bmj.m441
openalex publication_date 2020/03/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01
Abstract
Clinical prediction models aim to predict outcomes in individuals, to inform diagnosis or prognosis in healthcare. Hundreds of prediction models are published in the medical literature each year, yet many are developed using a dataset that is too small for the total number of participants or outcome events. This leads to inaccurate predictions and consequently incorrect healthcare decisions for some individuals. In this article, the authors provide guidance on how to calculate the sample size required to develop a clinical prediction model.
Citations
Cited by
- Differences in Mortality Between Treatment-Naive and Treatment-Discontinuing Hospitalized Individuals With Advanced HIV Disease: A Comparative Retrospective Study From Mexico City
- CBCT-Based Clinico-Radiomic Nomogram Predicting Preoperative Mandibular Third Molar Difficulty: Development/Validation
- A comparison of hyperparameter tuning procedures for clinical prediction models: A simulation study
- Prognostic models predicting clinical outcomes in patients diagnosed with visceral leishmaniasis: a systematic review
- Development and internal validation of a diagnostic prediction model for life-threatening events in callers with shortness of breath: a cross-sectional study in out-of-hours primary care
- Rapid prediction of cerebral edema on CT scan after traumatic brain injury
- Association of the systemic immune inflammation index with failure after core decompression for osteonecrosis of the femoral head: a prospective time-to-event analysis
- The harm of class imbalance corrections for risk prediction models: illustration and simulation using logistic regression
- Uncertainty of risk estimates from clinical prediction models: rationale, challenges, and approaches
- TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods
- Comparative evaluation of low- and high-field 1H NMR for vegetable oil authentication: classification, quantification, and routine applicability
- Time of coronary revascularization: methodology of a mediation analysis study
- Clinical prediction models and the multiverse of madness
- Improving embryo ploidy prediction: a machine learning approach using morphokinetic meta-variables and clinical data
- Minimum sample size for developing a multivariable prediction model using multinomial logistic regression
- Evaluating the sample size requirements of tree-based ensemble machine learning techniques for clinical risk prediction
- Surgical precision? Cutting through the hype of AI‐augmented peri‐operative risk prediction
- When the whole is greater than the sum of its parts: why machine learning and conventional statistics are complementary for predicting future health outcomes
- Predicting type 2 diabetes and testosterone effects in high-risk Australian men: development and external validation of a 2-year risk model
- PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods
- Prediction models for diagnosis and prognosis of covid-19: systematic review and critical appraisal
- A tutorial on individualized treatment effect prediction from randomized trials with a binary endpoint
- Association between periodontal diseases and cardiovascular diseases, diabetes and respiratory diseases: Consensus report of the Joint Workshop by the European Federation of Periodontology (EFP) and the European arm of the World Organization of Family Doctors (WONCA Europe)
- A decomposition of Fisher's information to inform sample size for developing or updating fair and precise clinical prediction models -- Part 3: continuous outcomes
- A novel scoring algorithm for chest pain can effectively support the diagnosis of acute coronary syndrome in prehospital settings: a cross-sectional study
- Three-year cardiovascular risk prediction among people who use cocaine or methamphetamine
- Machine learning vs. traditional methods for predicting postoperative cardiac complications after non‐cardiac surgery: a systematic review and Bayesian network meta‐analysis
- Development of a risk model for low oocyte retrieval in first-cycle IVF patients with diminished ovarian reserve: a retrospective single-center study
- Dynamic Updating of Psychosis Prediction Models in Individuals at Ultra-High Risk of Psychosis
- Comparing methods for handling missing data in electronic health records for dynamic risk prediction of central-line associated bloodstream infection
- The fundamental problem of risk prediction for individuals: health AI, uncertainty, and personalized medicine
- powerROC: An Interactive Web Tool for Sample Size Calculation in Assessing Models' Discriminative Abilities
- Imputation and missing indicators for handling missing data in the development and deployment of clinical prediction models: A simulation study
- Causal analyses of existing databases: the importance of understanding what can be achieved with your data before analysis (commentary on Hernán)
- Person-specific and pooled prediction models for binge eating, alcohol use and binge drinking in bulimia nervosa and alcohol use disorder
- Peer review of clinical and translational research manuscripts: Perspectives from statistical collaborators
- Adaptive Gaussian Process Search for Simulation-Based Sample Size Estimation in Clinical Prediction Models: Validation of the pmsims R Package
- A Nomogram Predicts the Need for Internal Iliac Vein Dissection During Renal Transplantation: A Multicenter Collaborative Study
Related