Related Experiment Video
Updated: Aug 24, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Flexible propensity score estimation strategies for clustered data in observational studies
Ting-Hsuan Chang1, Trang Quynh Nguyen2, Youjin Lee3
1Department of Epidemiology, Johns Hopkins Bloomberg School of Public Health, Baltimore, Maryland, USA.
Nonparametric machine learning methods may outperform logistic regression for propensity score estimation in large clustered studies, but may be vulnerable to unmeasured confounding in smaller clusters.
Area of Science:
- Epidemiology
- Biostatistics
- Machine Learning
Background:
- Nonparametric machine learning often shows superior performance to logistic regression for propensity score estimation.
- The effectiveness of nonparametric methods in clustered settings, particularly with unmeasured cluster-level confounding, remains unclear.
Purpose of the Study:
- To compare the performance of logistic regression, Bayesian additive regression trees, and generalized boosted modeling for propensity score weighting in clustered settings.
- To evaluate the impact of sample size, cluster size, and unmeasured cluster-level confounding on these methods.
Main Methods:
- Simulated data from three hypothetical observational studies with varying sample and cluster sizes.
- Included cluster indicators or random intercepts to account for clustering.
- Generated confounders at individual and cluster levels, including an unobserved cluster-level confounder.
- Assessed performance based on covariate balance, bias reduction, and confidence interval coverage.
- Applied methods to the National Longitudinal Study of Adolescent to Adult Health data.
Main Results:
- Nonparametric methods demonstrated better covariate balance, bias reduction, and confidence interval coverage in large clustered settings, irrespective of model complexity.
- In small sample or cluster sizes, nonparametric methods showed increased vulnerability to unmeasured cluster-level confounding.
- Multilevel logistic regression may be a more robust alternative in smaller, clustered settings with potential unmeasured confounding.
Conclusions:
- Nonparametric propensity score estimation offers advantages in large clustered datasets but requires careful consideration in smaller clusters due to potential unmeasured confounding.
- The choice of method should account for sample size, cluster size, and the potential for unmeasured cluster-level confounders.
More Related Videos
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Study Design in Statistics
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...

