Related Experiment Video
Updated: Aug 5, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Understanding differences between what alternate propensity score methods estimate
Anirban Basu1, Aig Unuigbe2, Cristina Masseria3
1The Comparative Health Outcomes, Policy, and Economics (CHOICE) Institute, University of Washington, Seattle.
Abstract:
BACKGROUND: Many approaches to propensity score methods are used in the applied health economics and outcomes research literature. Often this creates confusion when different approaches produce different results for the same data. OBJECTIVE: To present a conceptual overview based on a potential outcomes framework to demonstrate how more than 1 mean treatment effect parameter can be estimated using the propensity score methods and how the selection of appropriate methods should align with the scientific questions. METHODS: We highlight that more than 1 mean treatment effect parameter can be estimated using the propensity score methods. Using the potential outcomes framework and alternate data-generating processes, we discuss under what assumptions different mean treatment effect parameter estimates are supposed to vary. We tie these discussions with propensity score methods to show that different approaches may estimate different parameters. We illustrate these methods using a case study of the comparative effectiveness of apixaban vs warfarin on the likelihood of stroke among patients with a prior diagnosis of atrial fibrillation. RESULTS: Different mean treatment effect parameters take on different values when treatment effects are heterogeneous. We show that traditional propensity score approaches, such as blocking, weighting, matching, or doubly robust, can estimate different mean treatment effect parameters. Therefore, they may not produce the same results even when applied to the same data using the same covariates. We found significant differences in our case study estimates of mean treatment effect parameters. Still, once a mean treatment effect parameter is targeted, estimates across different methods are not different. This highlights the importance of first selecting the target parameter for analysis by aligning the interpretation of the target parameter with the scientific questions and then selecting the specific method to estimate this target parameter. CONCLUSIONS: We present a conceptual overview of propensity score methods in health economics and outcomes research from a potential outcomes framework. We hope these discussions will help applied researchers choose appropriate propensity score approaches for their analysis. DISCLOSURES: Dr Unuigbe's time was supported through an unrestricted postdoctoral fellowship from Pfizer to the University of Washington, Seattle.
Related Concept Videos
Comparing the Survival Analysis of Two or More Groups
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Variation
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
Null and Alternative Hypotheses
The null hypothesis, denoted by H0 is a statement of no difference between the variables—they are not related. This can often be considered the status quo. As a result if you cannot accept the null, it requires some action.
The alternative hypothesis, denoted by H1 or Ha, is a claim about the...
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Group Design

