Related Experiment Videos
Sample size calculation for training ensemble machine learning models on health data
Nicholas Mitsakakis1, Dan Liu1,2, Thomas Walters3
1CHEO Research Institute, Ottawa, ON, Canada.
Abstract:
Health research studies often suffer from small sample sizes, and training machine learning (ML) models requires large datasets. There is a dearth of literature on determining the adequate sample size for using ML models. We developed an empirically derived sample size calculator for ensemble ML models: random forests and two gradient-boosted decision trees (light gradient boosting machine [LGBM] and extreme gradient boosting [XGBoost]). This predicts the sample size required to achieve a pre-defined level of prognostic performance with a certain probability. Prognostic performance is defined as the sample area under the ROC curve (ROC-AUC) relative to the optimal model trained on the full (population) dataset. Our calculator's accuracy was compared to three common heuristics and a statistical approach to sample size calculation. For example, the median relative error sample size prediction was 25% to achieve 85% of the optimal performance with 90% certainty for LGBM. Our model has significantly better accuracy than other methods for tree-based ensemble ML models.
Related Concept Videos
Sample Size Calculation
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
Bootstrapping
Estimating Population Standard Deviation
Mechanistic Models: Compartment Models in Individual and Population Analysis
Estimating Population Mean with Known Standard Deviation
The confidence interval estimate will have the form as follows:
(point estimate - error bound, point estimate + error bound)
The...
Estimating Population Mean with Unknown Standard Deviation
William S. Gosset (1876–1937) of the Guinness...