Privacy-protecting multivariable-adjusted distributed regression analysis for multi-center pediatric study

Sengwee Toh1, Sheryl L Rifas-Shiman2, Pi-I D Lin2

  • 1Therapeutics Research and Infectious Disease Epidemiology Group, Department of Population Medicine, Harvard Pilgrim Health Care Institute, Harvard Medical School, Boston, MA, USA. darren_toh@harvardpilgrim.org.

Pediatric Research
|October 3, 2019
PubMed

Insights

Distributed regression, a privacy-preserving method, was validated in a large multi-center pediatric study. This approach enables multi-center research without sharing sensitive individual-level data, crucial for vulnerable populations.

Area of Science:

  • Pediatric Health Research
  • Biostatistics
  • Data Privacy

Background:

  • Privacy-preserving analytic methods are vital for vulnerable populations like children.
  • Distributed regression has not been previously tested in multi-center pediatric studies.

Purpose of the Study:

  • To assess the feasibility and validity of distributed linear regression in a multi-center pediatric study.
  • To compare distributed regression with conventional pooled individual-level data analysis.

Main Methods:

  • Utilized electronic health data from 34 healthcare institutions (PCORnet).
  • Fitted 12 multivariable-adjusted linear regression models assessing antibiotic use and BMI z-score.
  • Compared results from pooled individual-level data analysis and distributed regression using summary-level data.

Main Results:

  • Distributed linear regression and pooled individual-level analyses yielded nearly identical parameter estimates and standard errors.
  • The maximum difference in parameter estimates or standard errors was extremely small (4.4833 × 10⁻¹⁰).

Conclusions:

  • Empirically demonstrated the feasibility and validity of distributed linear regression in a large multi-center pediatric study.
  • This privacy-preserving method can facilitate multi-center pediatric research where data sharing is difficult.
Abstract

Related Concept Videos

Model Approaches for Pharmacokinetic Data: Distributed Parameter Models01:06

Model Approaches for Pharmacokinetic Data: Distributed Parameter Models

Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
218
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis00:59

Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis

Noncompartmental analyses offer an alternative method for describing drug pharmacokinetics without relying on a specific compartmental model. In this approach, the drug's pharmacokinetics are assumed to be linear, with the terminal phase log-linear. This assumption allows for simplified analysis and interpretation of the drug's behavior in the body.
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
288
Study Design in Statistics01:15

Study Design in Statistics

A study design is a set of techniques that allow a researcher to collect and analyze data from different variables defined for a specific research problem. Statistics is commonly for effective study design and more robust experiments,
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
9.9K
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
858
Multiple Regression01:25

Multiple Regression

Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K
Statistical Software for Data Analysis and Clinical Trials01:12

Statistical Software for Data Analysis and Clinical Trials

Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
1.4K