Related Experiment Video
Updated: Feb 28, 2026

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Evaluating the sampling effect of propensity score matching for reducing selection bias in medical data
Minji Roh1, Sujin Yum1, Gihun Joo2
1Interdisciplinary Graduate Program in Medical Bigdata Convergence, Kangwon National University, Chuncheon, Republic of Korea.
Background:
In real-world medical data, selection bias can significantly impact the performance of machine learning models, potentially leading to distorted outcomes. However, research aimed at mitigating selection bias remains relatively limited.
Methods:
In this study, we evaluate the effectiveness of Propensity Score Matching (PSM) in reducing selection bias and assessing its impact on classification performance in imbalanced medical data. Specifically, we apply PSM alongside five undersampling, three oversampling, and three hybrid sampling techniques to three medical datasets: rapidly progressive dementia prediction (ADNI, n = 628, events = 51), hypothyroidism prediction (UCI, n = 3,772, events = 3,481), and cardiovascular disease prediction (Kaggle, n = 253,680, events = 23,893), each exhibiting varying degrees of demographic selection bias. We train and compare six classification models to assess the impact of each resampling technique on model performance. The magnitude of selection bias is quantified using the standardized mean difference (SMD), while model performance is assessed using the Area Under the Receiver Operating Characteristic Curve (AUROC), the Area Under the Precision-Recall Curve (AUPRC), accuracy, precision, recall, F1-score, specificity, calibration curves, Brier score, and decision curve analysis.
Results:
The results indicate that PSM reduces SMD within the dataset, maintains stable classification performance, and enhances the internal validity of the model under conditions of limited or moderate demographic imbalance.
Conclusion:
These advantages suggest its potential for improving model reliability and facilitating better generalization to external datasets in real-world medical applications. However, in datasets with extreme selection bias or when overly restrictive matching is applied, PSM can degrade model performance, underscoring the importance of choosing strategies that account for dataset characteristics.
More Related Videos
03:05Influence of Emotional Factors on the Efficacy of Acupuncture Treatment for Overweight Complicated with Hyperlipidemia: A Retrospective Cohort Study
Published on: November 21, 2025
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Related Concept Videos
Kaplan-Meier Approach
Comparing the Survival Analysis of Two or More Groups
Sign Test for Matched Pairs
To conduct the sign test, we first calculate the differences in...
Bias in Epidemiological Studies
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Regression Toward the Mean