Related Experiment Video
Updated: Jul 26, 2025

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Avoiding Biased Clinical Machine Learning Model Performance Estimates in the Presence of Label Selection
Conor K Corbin1,2, Michael Baiocchi3, Jonathan H Chen2,4,5
1Department of Biomedical Data Science, Stanford, California, USA.
Estimating clinical machine learning model performance requires careful consideration of the deployment population. This study shows that label selection biases performance metrics, but causal inference weighting estimators can recover accurate estimates for the full population.
Area of Science:
- Machine Learning
- Clinical Informatics
- Causal Inference
Background:
- Evaluating clinical machine learning models requires understanding the deployment population.
- Label selection, where observed patients are a subset of the deployment population, can lead to misleading performance estimates.
- Standard metrics may not accurately reflect real-world performance due to biased label selection.
Purpose of the Study:
- To describe classes of label selection and simulate scenarios to assess bias in machine learning performance metrics.
- To investigate how label selection mechanisms affect model discrimination and calibration.
- To propose and evaluate methods for obtaining accurate performance estimates in deployed clinical models.
Main Methods:
- Simulated five causally distinct scenarios of label selection.
- Assessed bias in commonly reported binary machine learning performance metrics.
- Applied traditional weighting estimators from causal inference.
- Trained machine learning models to flag low-yield laboratory diagnostics.
- Proposed an altered deployment procedure combining randomization and weighted estimates.
Main Results:
- Selection affected by observed features can mislead discrimination estimates.
- Selection affected by labels can mislead calibration estimates.
- Weighting estimators, when properly specified, recover full population estimates.
- Naive AUROC estimates on the observed population undershot actual performance by up to 20% in a real-world task.
- The proposed deployment procedure recovered true model performance.
Conclusions:
- Label selection poses a significant challenge to accurate clinical machine learning model evaluation.
- Causal inference methods, particularly weighting estimators, are crucial for correcting performance bias.
- Misleading performance estimates can lead to premature termination of beneficial clinical tools.
- An altered deployment strategy combining randomization and weighting is effective in recovering true model performance.
Related Concept Videos
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Bias in Epidemiological Studies
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...

