Related Experiment Video
Updated: Oct 17, 2025

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Developing clinical prediction models when adhering to minimum sample size recommendations: The importance of
Glen P Martin1, Richard D Riley2, Gary S Collins3
1Division of Informatics, Imaging and Data Science, Faculty of Biology, Medicine and Health, 5292University of Manchester, Manchester Academic Health Science Centre, UK.
Minimum sample size formulas help prevent overfitting in clinical prediction models. Penalisation methods further reduce average overfitting but increase performance variability, necessitating careful examination during model development.
Area of Science:
- Statistics
- Biostatistics
- Clinical Epidemiology
Background:
- Minimum sample size formulas (e.g., Riley et al.) aim to prevent overfitting in clinical prediction model development.
- While effective on average, the variability of overfitting at recommended sample sizes remains unclear.
Purpose of the Study:
- To investigate the variability of overfitting in logistic regression clinical prediction models using simulation and empirical data.
- To compare unpenalised maximum likelihood estimation with post-estimation shrinkage or penalisation methods.
Main Methods:
- Simulation study and empirical example.
- Development of logistic regression models using unpenalised maximum likelihood estimation.
- Application of various post-estimation shrinkage or penalisation techniques.
Main Results:
- Mean calibration slopes were near one for all methods, indicating good average calibration.
- Penalisation methods reduced average overfitting compared to unpenalised methods.
- Penalisation methods exhibited higher variability in external predictive performance.
Conclusions:
- Penalisation methods are recommended alongside minimum sample size requirements to further mitigate overfitting in clinical prediction models.
- Examining variability in predictive performance and tuning parameters is crucial for robust model development and assures performance in new individuals.
Related Concept Videos
Bootstrapping
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Survival Tree
Building a Survival Tree
Constructing a...
Sample Size Calculation
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.

