Related Experiment Videos
Recent criteria and simple rules agreed often on required sample size for developed clinical prediction models
Ewout Steyerberg1, Toby Hackmann2, Ben van Calster3
1Julius Center for Health Sciences and Primary Care, University Medical Center Utrecht, Utrecht, the Netherlands; Department of Biomedical Data Sciences, Leiden University Medical Center, Leiden, the Netherlands.
Objective:
Adequate sample size is essential to the development of new prediction models with binary outcomes. We aim to relate recent approaches to traditional rules of thumb ('simple rules') and assess agreement and differences in judging adequacy of sample size for developed prediction models.
Study Design And Setting:
A well-known simple rule considers the events per variable (EPV), or more specifically, events per predictor parameter p (EPP, e.g. 'EPP>10', or 'EPP>20') to limit overfitting to small data sets. Another simple rule is to require a minimum absolute number of events (E) for reliable estimation of the overall event rate (e.g. 'E>100'). Recent criteria for regression-based prediction models consider: 1) limited overfitting in predictor effect estimates ('global shrinkage ≥ 0.9'); 2) small optimism in Nagelkerke's R2 ('δ(R2) ≤ 0.05'); and 3) precise estimation of the overall event rate (margin of error <0.05). We use statistical theory to compare simple rules to these three recent criteria. Furthermore, we compare sample size assessments for 299 published COVID-19 prediction models.
Results:
At a 10% event rate, criterion 1 (limited overfitting) corresponded to the classic EPP>10 rule for an Area under the ROC curve (c) of 0.767, and EPP>20 for c= 0.694. Criterion 2 implied EPP>4 (largely irrespective of c) and criterion 3 was equivalent to E>14 at a 10% event rate. We classified largely the same COVID-19 prediction models as adequate or inadequate for sample size according to recent criteria (maximum of three) and a simple combination rule ('E=100 plus 10*p', or max('E=100, 10*p)).
Conclusion:
Simple rules have direct relations with recent criteria for sample size calculations to limit overfitting and to guarantee reliability of predictions. Recent criteria are essential to inform prediction modeling efforts, while simple rules may often be sufficient to help appraise the quality of already developed regression-based prediction models.
Related Concept Videos
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Receiver Operating Characteristic Plot
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Sample Size Calculation
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
Hazard Ratio
For example, in a clinical trial evaluating a...
Study Design in Statistics
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...