Related Experiment Video
Updated: May 30, 2025

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
The Harms of Class Imbalance Corrections for Machine Learning Based Prediction Models: A Simulation Study
Alex Carriero1, Kim Luijken1, Anne de Hond1
1Julius Center for Health Sciences and Primary Care, University Medical Center Utrecht, Utrecht, The Netherlands.
Correcting for class imbalance in machine learning models may harm calibration. Models without imbalance correction consistently showed better or equal calibration, avoiding risk over-estimation in clinical prediction.
Area of Science:
- Machine Learning
- Biostatistics
- Clinical Informatics
Background:
- Risk prediction models are vital for clinical decisions, requiring accurate calibration.
- Healthcare data often exhibit class imbalance, leading researchers to apply corrections.
- The impact of these imbalance corrections on model calibration remains unclear.
Purpose of the Study:
- To investigate the effect of class imbalance corrections on the calibration of machine learning models.
- To compare calibration performance between models with and without imbalance correction across various scenarios.
Main Methods:
- Utilized extensive Monte Carlo simulations to assess out-of-sample predictive performance.
- Evaluated machine learning algorithms under different data-generating conditions (sample size, predictors, event fraction).
- Illustrated findings with a case study using MIMIC-III data.
Main Results:
- Models developed without class imbalance correction consistently demonstrated superior or equivalent calibration.
- Imbalance correction led to miscalibration, characterized by risk over-estimation.
- Re-calibration did not always resolve the miscalibration introduced by imbalance correction.
Conclusions:
- Class imbalance correction is not universally required for clinical prediction models.
- Applying imbalance correction can potentially impair model calibration and reliability.
- Prioritize model calibration over automatic imbalance correction for individual risk estimation.
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Regression Toward the Mean
Mechanistic Models: Compartment Models in Individual and Population Analysis
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Stereotype Content Model

