Related Experiment Video
Updated: Jul 4, 2026

Inverse Probability of Treatment Weighting (Propensity Score) using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Weighting regressions by propensity scores.
David A Freedman1, Richard A Berk
1University of California, Berkeley, Berkeley, CA, USA. freedman@stat.berkeley.edu
Propensity score weighting can reduce bias in regression analyses but may increase random error and bias standard errors. A well-specified causal model is often preferable to weighting for accurate parameter estimation.
Area of Science:
- Epidemiology
- Biostatistics
- Causal Inference
Background:
- Propensity score weighting is a common method to address confounding in observational studies.
- Weighting aims to create pseudo-populations where treatment assignment is independent of observed covariates.
Purpose of the Study:
- To evaluate the impact of propensity score weighting on bias and precision of causal parameter estimates.
- To compare weighting methods with direct causal modeling approaches.
Main Methods:
- The study theoretically analyzes the effects of propensity score weighting on regression estimates.
- It considers scenarios with both correctly and incorrectly specified causal models.
Main Results:
- Weighting can reduce bias but often inflates random error and biases standard error estimates downwards.
- In certain situations, propensity score weighting may exacerbate bias in causal parameters.
- Directly fitting a well-specified causal model generally yields more reliable estimates than weighting.
Conclusions:
- While propensity score weighting can be useful, its limitations regarding error inflation and potential bias must be considered.
- Investigators with a strong causal model should prioritize fitting the model directly.
- Weighting may offer some benefits when causal models are misspecified, but significant challenges remain.
Related Concept Videos
Regression Toward the Mean
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Weighted Mean
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
