Related Experiment Video
Updated: May 17, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
How Effective Are Machine Learning and Doubly Robust Estimators in Incorporating High-Dimensional Proxies to Reduce
Mohammad Ehsanul Karim1,2, Yang Lei3
1School of Population and Public Health, University of British Columbia, Vancouver, British Columbia, Canada.
High-dimensional proxies improve confounding adjustment in observational studies. Standard methods with proxies are robust, while complex machine learning in TMLE may reduce coverage, necessitating careful configuration.
Area of Science:
- Epidemiology
- Biostatistics
- Machine Learning in Health Research
Background:
- Residual confounding is a major challenge in high-dimensional observational studies.
- High-dimensional proxy adjustment methods, like hdPS, use proxies for unmeasured confounders.
- Machine learning and doubly robust estimators have been integrated into hdPS extensions, but their comparative performance is unclear.
Purpose of the Study:
- To evaluate the performance of standard methods, super learner (SL), targeted maximum likelihood estimation (TMLE), and double cross-fit TMLE (DC-TMLE) in confounding adjustment.
- To compare these methods under varying exposure and outcome prevalence using different machine learning learner configurations.
- To assess the impact of high-dimensional proxies and learner complexity on bias, coverage, and variability.
Main Methods:
- Plasmode simulations were conducted to assess method performance.
- Evaluated standard methods, SL, TMLE, and DC-TMLE.
- Compared performance across three learner library sizes: 1, 3, and 4 learners (including logistic regression, MARS, LASSO, and XGBoost).
Main Results:
- Methods without proxies showed the highest bias and lowest coverage.
- Standard methods incorporating high-dimensional proxies demonstrated robust performance with low bias and good coverage.
- TMLE and DC-TMLE reduced bias but had worse coverage, especially with larger, complex learner libraries.
- DC-TMLE underperformed in high-dimensional settings with non-Donsker learners, indicating instability.
Conclusions:
- High-dimensional proxies are crucial for effective confounding adjustment in standard methods.
- Tailoring machine learning learner configurations in SL and TMLE is essential for reliable confounding adjustment.
- Careful selection of learners is vital to avoid instability and ensure accurate results in complex observational studies.
Related Concept Videos
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding in Epidemiological Studies
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Regression Toward the Mean
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...

