Related Experiment Video
Updated: Jul 18, 2025

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Sample size requirements are not being considered in studies developing prediction models for binary outcomes: a
Paula Dhiman1, Jie Ma2, Cathy Qi3
1Centre for Statistics in Medicine, Nuffield Department of Orthopaedics, Rheumatology and Musculoskeletal Sciences, University of Oxford, Oxford, OX3 7LD, UK. paula.dhiman@csm.ox.ac.uk.
Most clinical prediction models lack proper sample size calculations, leading to insufficient data for accurate risk estimation. Researchers must justify, perform, and report sample size calculations for reliable model development.
Area of Science:
- Clinical prediction modeling
- Biostatistics
- Health research methodology
Background:
- Appropriate sample size is crucial for developing reliable clinical prediction models.
- This study reviews sample size considerations in models predicting binary outcomes.
Approach:
- Searched PubMed for studies (July 2020) developing prediction models for binary outcomes.
- Reviewed sample size calculations, comparing used vs. required sizes for risk estimation and overfitting minimization.
- Calculated minimum sample sizes and assessed adherence to the 10 events per variable (EPV) rule.
Key Points:
- Only 8% of 119 studies justified sample size.
- 73% used insufficient sample sizes to estimate overall risk and minimize overfitting.
- 75% did not meet the 10 EPV criteria, impacting model performance (median c-statistic 0.80).
Conclusions:
- Clinical prediction models are frequently developed without adequate sample size calculations, compromising risk estimation accuracy.
- Researchers are urged to justify, perform, and report sample size calculations for robust prediction model development.
More Related Videos
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Related Concept Videos
Sample Size Calculation
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
Survival Tree
Building a Survival Tree
Constructing a...
Odds Ratio
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Choosing Between z and t Distribution
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.