Related Experiment Video
Updated: Aug 20, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
An overview of propensity score matching methods for clustered data
Benjamin Langworthy1,2, Yujie Wu1, Molin Wang1,2
1Department of Biostatistics, Harvard T.H. Chan School of Public Health, Boston, MA, USA.
Propensity score matching in clustered data can control for measured and unmeasured confounders. Machine learning methods may underperform with many clusters, but alternative strategies exist for robust causal inference.
Area of Science:
- Epidemiology
- Biostatistics
- Observational Studies
Background:
- Propensity score matching (PSM) is vital for causal inference in observational studies.
- Clustered data present unique challenges for standard PSM methods.
- Existing methods struggle to balance measured and unmeasured confounders effectively in clustered settings.
Purpose of the Study:
- To provide an overview of PSM techniques for clustered data.
- To explore the utility of machine learning for estimating propensity scores in clustered data.
- To compare the performance of various PSM methods for clustered data.
Main Methods:
- Overview of PSM for clustered observational data.
- Application of generalized boosted models for propensity score estimation.
- Simulation studies comparing PSM methods under different clustering scenarios.
- Illustrative analysis using aspirin's effect on hearing deterioration.
Main Results:
- PSM can account for measured and unmeasured cluster-level confounders.
- Machine learning methods (e.g., GBM) for propensity scores can degrade with high clustering.
- Alternative strategies combining PSM with fixed-effects models can mitigate performance issues.
- Simulation results highlight the performance variations of different PSM approaches.
Conclusions:
- PSM is adaptable for clustered data, offering control over various confounders.
- Careful consideration of methods is needed when using machine learning with highly clustered data.
- Hybrid approaches may enhance causal effect estimation in complex observational studies.
- The study provides practical insights for analyzing clustered observational data.
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Randomized Experiments
Simple randomization
Simple...
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Comparing the Survival Analysis of Two or More Groups
Wilcoxon Signed-Ranks Test for Matched Pairs
Sign Test for Matched Pairs
To conduct the sign test, we first calculate the differences in...

