Related Experiment Video
Updated: Jan 31, 2026

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Combining distributed regression and propensity scores: a doubly privacy-protecting analytic method for multicenter
Sengwee Toh1, Robert Wellman2, R Yates Coley2
1Department of Population Medicine, Harvard Medical School and Harvard Pilgrim Health Care Institute, Boston, MA, USA, darren_toh@harvardpilgrim.org.
Distributed linear regression is a feasible and valid method for multicenter studies, enabling multivariable-adjusted analysis using only summary-level data. Combining this approach with propensity scores enhances privacy and analytical flexibility for continuous outcomes.
Area of Science:
- Biostatistics
- Health Informatics
- Epidemiology
Background:
- Sharing detailed individual-level data in multicenter studies presents significant privacy and logistical challenges.
- Analytic methods utilizing summary-level data offer a potential solution for privacy-preserving multivariable analyses.
Purpose of the Study:
- To assess the feasibility and validity of multivariable-adjusted distributed linear regression.
- To evaluate the combination of distributed linear regression with propensity scores in a large distributed data network.
Main Methods:
- Compared percent total weight loss 1-year post-surgery between Roux-en-Y gastric bypass and sleeve gastrectomy using data from 43,110 patients across 36 health systems.
- Employed distributed linear regression, which uses only summary-level data (sums of squares and cross products matrix), to fit three regression models.
- Adjusted for baseline variables using individual covariates, propensity score deciles, or both, comparing results to analyses using pooled individual-level data.
Main Results:
- Distributed linear regression yielded results identical to pooled individual-level data analyses across all variables and models.
- The maximum numerical difference in parameter estimates or standard errors was negligible (3×10-11).
Conclusions:
- Distributed linear regression is a feasible and valid method for analyzing continuous outcomes in multicenter studies.
- Integrating distributed regression with propensity score modeling improves privacy protection and analytical flexibility.
Related Concept Videos
Regression Toward the Mean
Analyte Adsorption and Distribution
Development of Analytical Methods
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Correlation and Regression
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:

