Related Experiment Video
Updated: Mar 3, 2026

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Sample size calculation to externally validate scoring systems based on logistic regression models
Antonio Palazón-Bru1, David Manuel Folgado-de la Rosa1, Ernesto Cortés-Castell2
1Department of Clinical Medicine, Miguel Hernández University, San Juan de Alicante, Alicante, Spain.
This study introduces an algorithm to determine sample size for validating binary logistic regression scoring systems. The developed algorithm found a sufficient sample size with fewer events than previously recommended.
Area of Science:
- Biostatistics
- Clinical Epidemiology
- Health Informatics
Background:
- Traditional recommendations suggest at least 100 events and 100 non-events for predictive model validation.
- Factors like discrimination, parameterization, and incidence can influence predictive model calibration.
- Scoring systems derived from binary logistic regression models are common predictive tools.
Purpose of the Study:
- To develop and present an algorithm for calculating the optimal sample size required for validating scoring systems based on binary logistic regression.
- To apply this novel algorithm to a real-world case study for practical demonstration.
Main Methods:
- The algorithm utilizes bootstrap sampling to assess key validation metrics.
- It calculates the area under the ROC curve (AUC) for discrimination.
- It evaluates observed event probabilities via smooth curves and the estimated calibration index (ECI) for calibration assessment.
Main Results:
- The case study application demonstrated the algorithm's efficacy.
- The algorithm determined an adequate sample size requiring only 69 events.
- This sample size is notably lower than the commonly cited benchmark of 100 events.
Conclusions:
- An effective algorithm is presented for determining the sample size needed to validate binary logistic regression-based scoring systems.
- This methodology offers a more precise approach to sample size calculation for model validation.
- The algorithm's applicability extends to similar validation scenarios in various research contexts.
More Related Videos
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
Related Concept Videos
Sample Size Calculation
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
Estimating Population Mean with Known Standard Deviation
The confidence interval estimate will have the form as follows:
(point estimate - error bound, point estimate +...
Range Rule of Thumb to Interpret Standard Deviation
For instance, the range rule of thumb can be used to find the tallest and the shortest student in a class, given the mean student height and standard deviation. If the mean student height is 1.6 m and the standard deviation, s is 0.05 m, the height...
Margin of Error
Bootstrapping
Chebyshev's Theorem to Interpret Standard Deviation