Related Experiment Video
Updated: Jul 28, 2025

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
Generating synthetic personal health data using conditional generative adversarial networks combining with
Chang Sun1, Johan van Soest2, Michel Dumontier1
1Institute of Data Science, Faculty of Science and Engineering, Maastricht University, Maastricht, The Netherlands; Department of Advanced Computing Sciences, Faculty of Science and Engineering, Maastricht University, Maastricht, The Netherlands.
Generating realistic synthetic health data is challenging. A new differentially private conditional Generative Adversarial Network (DP-CGANS) model addresses privacy and minority class data challenges, improving data utility and privacy balance.
Area of Science:
- Health Informatics
- Data Science
- Privacy-Preserving Technologies
Background:
- Personal health data is valuable but inaccessible due to privacy and legal constraints.
- Synthetic data offers a promising solution, but challenges remain in realism, privacy preservation, and handling imbalanced datasets.
- Existing methods struggle with minority class data simulation and capturing variable dependencies.
Purpose of the Study:
- To propose a novel differentially private conditional Generative Adversarial Network (DP-CGANS) model for generating realistic and privacy-preserving synthetic personal health data.
- To address the specific challenges of synthetic health data generation, including minority class representation and inter-variable dependency.
- To ensure a balance between data utility and patient privacy.
Main Methods:
- Developed a DP-CGANS model involving data transformation, sampling, conditioning, and network training.
- Distinguished and transformed categorical and continuous variables into latent space separately.
- Incorporated a conditional vector to represent minority classes and injected noise into gradients for differential privacy.
Main Results:
- The DP-CGANS model demonstrated superior performance in capturing variable dependencies compared to state-of-the-art models.
- Evaluations on socio-economic and real-world health datasets showed strong statistical similarity, machine learning utility, and privacy preservation.
- The model effectively balances data utility and privacy, even with imbalanced classes, abnormal distributions, and data sparsity.
Conclusions:
- DP-CGANS is an effective method for generating high-fidelity, privacy-preserving synthetic personal health data.
- The model successfully overcomes key challenges in synthetic health data generation, particularly for imbalanced and complex datasets.
- This approach enhances the accessibility of valuable health data for research while upholding stringent privacy standards.
More Related Videos
Related Concept Videos
Censoring Survival Data
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Genetic Material
What is Genetic Engineering?
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...

