Related Experiment Video
Updated: May 16, 2025

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Evaluating the sample size requirements of tree-based ensemble machine learning techniques for clinical risk
Oya Kalaycıoğlu1,2, Menelaos Pavlou2, Serhat E Akhanlı3
1Department of Biostatistics and Medical Informatics, Bolu Abant İzzet Baysal University, Bolu, Türkiye.
Sample size guidelines for clinical risk models using machine learning techniques (MLTs) are unclear. This study found that existing guidelines targeting mean absolute prediction error (MAPE) are insufficient for tree-based MLTs, but suitable for external validation of the C-statistic.
Area of Science:
- Biostatistics
- Machine Learning in Healthcare
- Clinical Prediction Modeling
Background:
- Machine learning techniques (MLTs) are widely adopted for clinical risk prediction.
- Clear sample size requirements for developing and validating MLT models are lacking.
- Existing guidelines often focus on logistic regression, not ensemble MLTs.
Purpose of the Study:
- To assess the applicability of sample size guidelines for logistic regression to tree-based ensemble MLTs (bagging, random forests, boosting).
- To evaluate MLT performance metrics (MAPE, C-statistic, Brier score, calibration) under various data-generating mechanisms and sample sizes.
- To determine appropriate sample size calculations for MLT development and external validation.
Main Methods:
- Simulations using two large cardiovascular datasets.
- Evaluation of MLTs (boosting, random forests, bagging) and logistic regression.
- Assessment across six data-generating mechanisms (DGMs) and varying sample sizes.
- Performance metrics included Mean Absolute Prediction Error (MAPE), C-statistic, Brier score, and calibration.
Main Results:
- Boosting models required 2-3 times the recommended sample size when the DGM matched the analysis model.
- Random forests and bagging failed to meet target MAPE even with a 12-fold sample size increase.
- Logistic regression and boosting required a 12-fold increase in sample size for a neutral DGM.
- Sample size guidelines for C-statistic precision were suitable for external validation of MLTs.
Conclusions:
- Existing sample size guidelines targeting MAPE are inadequate for tree-based ensemble MLTs.
- MLTs, particularly random forests and bagging, exhibit substantial sample size needs beyond current recommendations.
- Guidelines for C-statistic precision in external validation can inform sample size calculations for MLTs.
- Further research is needed to refine sample size recommendations for MLT model development.
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a...
Statistical Software for Data Analysis and Clinical Trials
Sample Size Calculation
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...

