Related Experiment Videos
Recent criteria and simple rules agreed often on required sample size for developed clinical prediction models
Ewout Steyerberg1, Toby Hackmann2, Ben van Calster3
1Julius Center for Health Sciences and Primary Care, University Medical Center Utrecht, Utrecht, The Netherlands; Department of Biomedical Data Sciences, Leiden University Medical Center, Leiden, The Netherlands.
Objectives:
Adequate sample size is essential to the development of new prediction models with binary outcomes. We aim to relate recent approaches to traditional rules of thumb ('simple rules') and assess agreement and differences in judging adequacy of sample size for developed prediction models.
Study Design And Setting:
A well-known simple rule considers the events per variable or, more specifically, events per predictor parameter p (EPP, eg, 'EPP > 10' or 'EPP > 20') to limit overfitting to small datasets. Another simple rule is to require a minimum absolute number of events (E) for reliable estimation of the overall event rate (eg, 'E > 100'). Recent criteria for regression-based prediction models consider: 1) limited overfitting in predictor effect estimates ('global shrinkage ≥0.9'); 2) small optimism in Nagelkerke's R2 ('δ(R2) ≤ 0.05'); and 3) precise estimation of the overall event rate (margin of error <0.05). We use statistical theory to compare simple rules to these three recent criteria. Furthermore, we compare sample size assessments for 299 published COVID-19 prediction models.
Results:
At a 10% event rate, criterion 1 (limited overfitting) corresponded to the classic EPP>10 rule for an area under the receiver operating characteristic curve (c) of 0.767 and EPP>20 for c = 0.694. Criterion 2 implied EPP>4 (largely irrespective of c), and criterion 3 was equivalent to E > 14 at a 10% event rate. We classified largely the same COVID-19 prediction models as adequate or inadequate for sample size according to recent criteria (maximum of three) and a simple combination rule ('E = 100 plus 10∗p', or max['E = 100, 10∗p']).
Conclusion:
Simple rules have direct relations with recent criteria for sample size calculations to limit overfitting and guarantee reliability of predictions. Recent criteria are essential to inform prediction modeling efforts, while simple rules may often be sufficient to help appraise the quality of already developed regression-based prediction models.
Plain Language Summary:
Adequate sample size is essential for reliable empirical research. We focus on the development of prediction models to provide reliable individualized estimates of the risk of an event. We compared traditional simple rules for deciding if a dataset is large enough to build a binary outcome prediction model (such as "events per predictor >10" or "total events >100") with more recent, refined statistical criteria. The newer criteria aim to guarantee (1) that predictions are not too extreme (global shrinkage ≥0.9) (2), that model performance optimism is small (decrease in R2 ≤ 0.05), and (3) that the overall event rate is precisely estimated. Using theory and examples from 299 published COVID-19 models, we found that the simple rules often map closely to the newer criteria. Overall, we conclude that while the new criteria are better for planning prospective studies and developing a prediction model, simple rules remain useful and are usually adequate for judging whether already developed prediction models had enough data.
Related Concept Videos
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Receiver Operating Characteristic Plot
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Sample Size Calculation
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
Hazard Ratio
For example, in a clinical trial evaluating a...
Study Design in Statistics
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...