Related Experiment Video
Updated: Sep 19, 2025

Methodology for Establishing a Community-Wide Life Laboratory for Capturing Unobtrusive and Continuous Remote Activity and Health Data
Published on: July 27, 2018
Generating synthetic electronic health record data: a methodological scoping review with benchmarking on phenotype
Xingran Chen1, Zhenke Wu1, Xu Shi1
1Department of Biostatistics, University of Michigan, Ann Arbor, MI 48109, United States.
This study benchmarks methods for generating synthetic Electronic Health Records (EHR) data, finding Generative Adversarial Network (GAN)-based approaches effective for data fidelity and utility. Rule-based methods offer superior privacy protection for EHR data generation.
Area of Science:
- Health Informatics
- Data Science
- Medical Data Generation
Background:
- Synthetic Electronic Health Records (EHR) data generation is crucial for research and development.
- Existing methods for synthetic EHR data generation vary in their effectiveness and application.
- A comprehensive review and benchmarking of these methods are needed to guide practitioners.
Purpose of the Study:
- To conduct a scoping review of synthetic EHR data generation approaches.
- To benchmark major synthetic data generation methods using open-source EHR datasets.
- To provide an open-source software tool and practical recommendations for synthetic EHR data generation.
Main Methods:
- A scoping review of three academic databases identified 48 relevant studies.
- Seven state-of-the-art methods and two baseline methods were implemented and benchmarked.
- Evaluation focused on data fidelity, downstream utility, privacy protection, and computational cost using MIMIC-III/IV datasets.
Main Results:
- Forty-eight studies were classified into five categories of synthetic data generation methods.
- Generative Adversarial Network (GAN)-based methods showed competitive performance in fidelity and utility.
- Rule-based methods demonstrated superior privacy protection, with similar trends observed on MIMIC-IV data.
Conclusions:
- Method selection depends on the prioritized evaluation metrics for specific use cases.
- A decision tree and an open-source Python package (SynthEHRella) are provided to aid practitioners.
- Future research should focus on enhancing data fidelity and privacy, and benchmarking longitudinal/conditional generation methods.
Related Concept Videos
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Methods of Documentation VII: EMR
Synthetic Biology
Golden rice
Golden rice is a genetically modified...

