Related Experiment Video
Updated: Jun 17, 2025

DetectSyn: A Rapid, Unbiased Fluorescent Method to Detect Changes in Synapse Density
Published on: July 22, 2022
Does Differentially Private Synthetic Data Lead to Synthetic Discoveries?
Ileana Montoya Perez1, Parisa Movahedi1, Valtteri Nieminen1
1Department of Computing, University of Turku, Turku, Finland.
Differential Privacy (DP) synthetic data can inflate statistical test errors, particularly Type I errors, when privacy budgets (epsilon) are low. A DP Smoothed Histogram method showed valid Type I errors but needed large datasets and higher privacy budgets for good power.
Area of Science:
- Biomedical Informatics
- Data Privacy
- Statistical Analysis
Background:
- Synthetic data generation is crucial for sharing sensitive biomedical information while preserving privacy.
- Differential Privacy (DP) is the standard for balancing data utility and individual privacy.
- Evaluating the trustworthiness of statistical findings from DP-synthetic data is essential.
Purpose of the Study:
- To assess the reliability of group difference discoveries from DP-synthetic data using independent sample tests.
- To quantify Type I (false discovery) and Type II (missed discovery) errors in statistical tests on DP-synthetic data.
- To understand the impact of different DP-synthetic data generation methods on statistical validity.
Main Methods:
- Evaluation of Mann-Whitney U, Student's t-test, chi-squared, and median tests.
- Generation of DP-synthetic data from real-world (prostate cancer, cardiovascular) and simulated datasets.
- Comparison of five DP-synthetic data generation methods, including histogram-based, MWEM, Private-PGM, and DP GAN.
Main Results:
- Many DP-synthetic data generation methods exhibited significantly inflated Type I errors at low privacy levels (epsilon <= 1).
- Low p-values in statistical tests may result from DP noise rather than true effects.
- A DP Smoothed Histogram method maintained valid Type I errors across privacy levels but required large datasets and epsilon >= 5 for acceptable Type II errors.
Conclusions:
- Caution is advised when interpreting statistical results from DP-synthetic data due to potential inflation of Type I errors.
- The choice of DP-synthetic data generation method critically impacts the validity and reliability of statistical analyses.
- Achieving both privacy and statistical accuracy requires careful consideration of dataset size, privacy budget, and generation methodology.
Related Concept Videos
Synthetic Biology
Golden rice
Golden rice is a genetically modified...
Censoring Survival Data
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Naturalistic Observations
Randomized Experiments
Simple randomization
Simple...
Correlation of Experimental Data
For example, a spherical particle moving through a viscous fluid experiences drag. Dimensional analysis shows that the drag force depends on the particle's diameter, velocity,...

