Related Experiment Video
Updated: Aug 18, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Exploring the utility of synthetic data to extract more value from sensitive health data assets: A focused example in
Amy Elise Braddon1, Suzanne Robinson1, Rosa Alati1
1School of Population Health, Curtin University, Western Australia, Perth, Australia.
Background:
Privacy, access and security concerns can hinder the availability of health data for research. The use of synthesised data in place of de-identified electronic health records (EHRs) presents an opportunity to conduct research while minimising privacy concerns.
Objectives:
To examine whether synthesised data can replicate two prenatal epidemiological associations: between prenatal smoking and lower birthweight, and between prenatal mood disorders and lower birthweight, using data synthesised from de-identified health administrative data collections.
Methods:
We generated two synthetic datasets, using parametric and non-parametric data generating methods, and examined the synthetic data for evidence of privacy concerns. Next, univariable and multivariable logistic regression was utilised to estimate the associations in both synthetic datasets, with results then compared to the real data.
Results:
Both synthesised datasets performed well in identifying the reduction in birthweight associated with prenatal smoking, while the non-parametric data underestimated the reduction in birthweight associated with prenatal mood disorders. Improbable relationships between some variables were identified in the parametric synthesised data, however, these can be addressed with simple rules during data synthesis. No duplicate rows (i.e., exact copies of de-identified data) were found in the parametric data, while only 0.6% of the rows in the non-parametric data were duplicated.
Conclusions:
Both synthesised datasets performed well in replicating the statistical properties of the original data while addressing privacy issues. Data synthesis methods provide an opportunity for researchers to utilise health data while managing privacy and security concerns.
More Related Videos
09:43Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering
Published on: November 22, 2019
09:33Visualizing Field Data Collection Procedures of Exposure and Biomarker Assessments for the Household Air Pollution Intervention Network Trial in India
Published on: December 23, 2022
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Statistical Software for Data Analysis and Clinical Trials
Analysis of Population Pharmacokinetic Data
Data Collection I