Related Experiment Video
Updated: Aug 4, 2026

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Substantial effective sample sizes were required for external validation studies of predictive logistic regression
Yvonne Vergouwe1, Ewout W Steyerberg, Marinus J C Eijkemans
1Department of Public Health, Erasmus MC, P.O. Box 1738, 3000 DR Rotterdam, The Netherlands. Y.Vergouwe@UMCUtrecht.nl
External validation of prediction models requires adequate sample sizes. This study suggests a minimum of 100 events and 100 nonevents for reliable detection of model performance differences.
Area of Science:
- Biostatistics
- Clinical Prediction Models
- Health Research Methodology
Background:
- Prediction model performance often declines in external validation datasets compared to development data.
- Assessing the reliability of prediction models in new populations is crucial for clinical utility.
- Determining optimal sample sizes for external validation is essential for robust model evaluation.
Purpose of the Study:
- To identify effective sample sizes, specifically the number of events, needed to detect meaningful differences in prediction model performance during external validation.
- To establish power requirements for detecting specific types of model invalidity.
Main Methods:
- Utilized logistic regression to predict benign tissue in residual masses for metastatic testicular cancer patients.
- Employed standard power calculations and Monte Carlo simulations to estimate required event numbers.
- Calculated sample sizes to detect specific model performance degradations with 80% power at a 5% significance level.
Main Results:
- Detecting overly optimistic probability predictions (1.5x odds scale) required 111 events in the validation sample.
- A decrease in discriminative ability (c-statistic from 0.83 to 0.73) necessitated 81 to 106 events, varying by scenario.
Conclusions:
- A minimum of 100 events and 100 nonevents is recommended for external validation samples.
- Substantially larger effective sample sizes may be necessary for specific hypothesis testing to achieve adequate statistical power.
Related Concept Videos
Sample Size Calculation
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
One-Way ANOVA: Unequal Sample Sizes
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...