Related Experiment Video
Updated: Nov 4, 2025

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Minimum sample size for external validation of a clinical prediction model with a binary outcome.
Richard D Riley1, Thomas P A Debray2, Gary S Collins3,4
1Centre for Prognosis Research, School of Medicine, Keele University, Staffordshire, UK.
Determining adequate sample size for external validation of prediction models is crucial. This study provides methods to calculate the minimum sample size needed for precise estimates of calibration, discrimination, and clinical utility in external validation studies.
Area of Science:
- Biostatistics
- Clinical Epidemiology
- Health Informatics
Background:
- External validation is essential for assessing prediction model generalizability.
- Current external validation studies often lack sufficient sample sizes, leading to unreliable performance estimates.
- Precise estimation of model performance is critical for clinical decision-making.
Purpose of the Study:
- To propose a method for determining the minimum sample size required for external validation studies of prediction models for binary outcomes.
- To enable precise estimation of calibration, discrimination, and clinical utility.
- To provide guidance on assessing the adequacy of existing datasets for external validation.
Main Methods:
- Development of closed-form and iterative solutions for sample size calculation.
- Inclusion of key parameters: target standard errors, event proportion, model calibration, and risk thresholds.
- Application of the method to a case study of mechanical heart valve failure prediction.
Main Results:
- Calculations indicate a need for at least 9835 participants (177 events) for precise calibration and discrimination estimates, often limited by the calibration slope.
- A sample size of 6443 participants (116 events) is required for precise net benefit estimation at an 8% risk threshold.
- The proposed methods can inform sample size adequacy for already collected datasets.
Conclusions:
- The proposed sample size calculation methods enhance the reliability of external validation studies for prediction models.
- Accurate sample size determination ensures precise performance estimates, supporting confident clinical application of prediction models.
- Availability of software code facilitates the implementation of these sample size calculations.
Related Concept Videos
Sample Size Calculation
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Margin of Error
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Receiver Operating Characteristic Plot
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...

