Related Experiment Video
Updated: Jan 6, 2026

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Privacy-protecting multivariable-adjusted distributed regression analysis for multi-center pediatric study
Sengwee Toh1, Sheryl L Rifas-Shiman2, Pi-I D Lin2
1Therapeutics Research and Infectious Disease Epidemiology Group, Department of Population Medicine, Harvard Pilgrim Health Care Institute, Harvard Medical School, Boston, MA, USA. darren_toh@harvardpilgrim.org.
Insights
Distributed regression, a privacy-preserving method, was validated in a large multi-center pediatric study. This approach enables multi-center research without sharing sensitive individual-level data, crucial for vulnerable populations.
Area of Science:
- Pediatric Health Research
- Biostatistics
- Data Privacy
Background:
- Privacy-preserving analytic methods are vital for vulnerable populations like children.
- Distributed regression has not been previously tested in multi-center pediatric studies.
Purpose of the Study:
- To assess the feasibility and validity of distributed linear regression in a multi-center pediatric study.
- To compare distributed regression with conventional pooled individual-level data analysis.
Main Methods:
- Utilized electronic health data from 34 healthcare institutions (PCORnet).
- Fitted 12 multivariable-adjusted linear regression models assessing antibiotic use and BMI z-score.
- Compared results from pooled individual-level data analysis and distributed regression using summary-level data.
Main Results:
- Distributed linear regression and pooled individual-level analyses yielded nearly identical parameter estimates and standard errors.
- The maximum difference in parameter estimates or standard errors was extremely small (4.4833 × 10⁻¹⁰).
Conclusions:
- Empirically demonstrated the feasibility and validity of distributed linear regression in a large multi-center pediatric study.
- This privacy-preserving method can facilitate multi-center pediatric research where data sharing is difficult.
Background:
Privacy-protecting analytic approaches without centralized pooling of individual-level data, such as distributed regression, are particularly important for vulnerable populations, such as children, but these methods have not yet been tested in multi-center pediatric studies.
Methods:
Using the electronic health data from 34 healthcare institutions in the National Patient-Centered Clinical Research Network (PCORnet), we fit 12 multivariable-adjusted linear regression models to assess the associations of antibiotic use <24 months of age with body mass index z-score at 48 to <72 months of age. We ran these models using pooled individual-level data and conventional multivariable-adjusted regression (reference method), as well as using the more privacy-protecting pooled summary-level intermediate statistics and distributed regression technique. We compared the results from these two methods.
Results:
Pooled individual-level and distributed linear regression analyses produced virtually identical parameter estimates and standard errors. Across all 12 models, the maximum difference in any of the parameter estimates or standard errors was 4.4833 × 10-10.
Conclusions:
We demonstrated empirically the feasibility and validity of distributed linear regression analysis using only summary-level information within a large multi-center study of children. This approach could enable expanded opportunities for multi-center pediatric research, especially when sharing of granular individual-level data is challenging.
Related Concept Videos
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
Study Design in Statistics
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Statistical Methods for Analyzing Epidemiological Data
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Statistical Software for Data Analysis and Clinical Trials

