Related Experiment Video
Updated: Jun 24, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
A comparison of random forest-based missing imputation methods for covariates in propensity score analysis.
Yongseok Lee1, Walter L Leite2
1Bureau of Economic and Business Research (BEBR), University of Florida.
This study evaluated imputation methods for missing data in propensity score analysis (PSA). Proximity imputation (PI) and missForest showed superior performance in reducing bias for treatment effect estimation.
Area of Science:
- Statistics
- Biostatistics
- Observational Research Methods
Background:
- Selection bias is a significant challenge in observational studies.
- Missing data in covariates complicates propensity score analysis (PSA).
- Effective handling of missing data is crucial for reliable PSA results.
Purpose of the Study:
- To evaluate multiple imputation methods using random forests for handling missing covariates in PSA.
- To compare the performance of multivariate imputation by chained equations-random forest (Caliber), proximity imputation (PI), and missForest.
- To assess the impact of imputation methods on the bias of average treatment effect estimates.
Main Methods:
- Monte Carlo simulations were employed to assess imputation method performance.
- Evaluated methods included Caliber, proximity imputation (PI), and missForest.
- Propensity score analysis was applied to a real-world dataset (Early Childhood Longitudinal Study).
Main Results:
- Proximity imputation (PI) and missForest demonstrated superior performance in reducing bias of the average treatment effect.
- The effectiveness of PI and missForest was consistent across different sample sizes and missing data mechanisms.
- The study provided a practical demonstration of these methods in evaluating educational interventions.
Conclusions:
- Proximity imputation (PI) and missForest are recommended for handling missing covariate data in propensity score analysis.
- These methods offer robust solutions for mitigating selection bias in observational research.
- Accurate estimation of treatment effects is enhanced by appropriate imputation techniques for missing data.
Related Concept Videos
Comparing the Survival Analysis of Two or More Groups
Randomized Experiments
Simple randomization
Simple...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Assumptions of Survival Analysis
Goodness-of-Fit Test
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...

