Related Experiment Video
Updated: Sep 17, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Machine Learning Versus Logistic Regression for Propensity Score Estimation: A Benchmark Trial Emulation Against the
Kaicheng Wang1,2, Lindsey A Rosman2, Haidong Lu3
1Yale Center for Analytical Sciences, Department of Biostatistics, Yale School of Public Health, New Haven, Connecticut, USA.
Machine learning propensity scores may introduce bias in heart failure studies. Traditional logistic regression with expert-guided confounder selection better estimated sacubitril/valsartan effectiveness compared to ML methods.
Area of Science:
- Cardiovascular Medicine
- Health Informatics
- Biostatistics
Background:
- Machine learning (ML) is increasingly used for propensity score estimation to improve covariate balance and reduce bias.
- The validity of ML in selecting appropriate confounders for causal inference remains controversial.
- Heart failure management often involves comparing drug effectiveness using real-world data.
Purpose of the Study:
- To estimate the effectiveness of sacubitril/valsartan versus traditional therapies on all-cause mortality in heart failure patients.
- To compare propensity score estimation using traditional logistic regression versus ML approaches.
- To benchmark real-world evidence (RWE) findings against the PARADIGM-HF randomized controlled trial.
Main Methods:
- Retrospective cohort study of U.S. Department of Veterans Affairs heart failure patients (2016-2020) with implantable cardioverter defibrillators.
- Propensity scores were estimated using logistic regression with *a priori* confounder selection and ML-based methods (generalized boosting models).
- Effectiveness was measured by all-cause mortality, hazard ratios (HR), and risk ratios (RR), compared to the PARADIGM-HF trial.
Main Results:
- Logistic regression with *a priori* confounder selection yielded results (HR=0.93, 95% CI 0.61-1.42) closely aligning with the PARADIGM-HF trial (HR=0.81, 95% CI 0.61-1.06).
- ML-based propensity scores, particularly with data-driven confounder selection, did not outperform logistic regression and potentially amplified bias (HR=0.63, 95% CI 0.31-1.30).
- ML methods may introduce overadjustment bias in complex, high-dimensional datasets.
Conclusions:
- Subject-matter expertise is crucial for confounder selection in causal inference with real-world data.
- Traditional logistic regression with careful confounder selection may be more reliable than ML approaches for estimating treatment effectiveness in this context.
- Over-reliance on data-driven methods without clinical validation can lead to biased estimates in cardiovascular research.
Related Concept Videos
Randomized Experiments
Simple randomization
Simple...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Study Design in Statistics
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Comparing the Survival Analysis of Two or More Groups
Regression Toward the Mean
Bias in Epidemiological Studies

