Related Experiment Video
Updated: May 31, 2026

Inverse Probability of Treatment Weighting (Propensity Score) using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Enhancing propensity score analysis with data missing not at random: Introducing dual-forest proximity imputation
Yongseok Lee1, Walter L Leite2
1Department of Human Development and Family Science, Purdue University, Hanley Hall (Room 356), 1202 Mitch Daniels Blvd, West Lafayette, IN, 47906, USA. lee5321@purdue.edu.
None:
Researchers using propensity score analysis (PSA) to estimate treatment effects using secondary data may have to handle data that are missing not at random (MNAR). Existing methods for PSA with MNAR data use logistic regression to model the missing data mechanisms, thus requiring manual specification of functional forms, and are difficult to implement with a large number of covariates. To overcome these limitations, this study proposes alternatives to existing methods by replacing logistic regression with a random forest. Also, it introduces the dual-forest proximity imputation method, which leverages two types of proximity matrices of random forest techniques and incorporates missingness pattern information in each matrix. Results from a Monte Carlo simulation show dual-forest proximity imputation's enhanced bias reduction with various types of MNAR mechanisms as compared to existing and alternative methods. A case study is also provided using data from the National Longitudinal Survey of Youth (Enders, 1979) (NLSY79).
Related Concept Videos
Assumptions of Survival Analysis
Truncation in Survival Analysis
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are observed.
Randomized Experiments
Simple randomization
Simple...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...
Optimal Foraging
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance, comparing...
