Related Experiment Video
Updated: Jan 14, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Fidelity-agnostic synthetic data generation improves utility while retaining privacy
Jim Achterberg1, Marcel Haas1, Bram van Dijk1
1Public Health and Primary Care (Health Campus The Hague), Leiden University Medical Center, Leiden, South-Holland, the Netherlands.
Abstract:
Synthetic data are a popular method to publish useful datasets in a privacy-aware manner, making them useful across a range of scientific domains involving human subjects. They are typically generated by sampling from algorithms that mimic the probability distribution of real datasets, thereby maximizing statistical similarity to real data. However, we argue and demonstrate that synthetic data need to be similar only in ways relevant to their intended use and may neglect any irrelevant information, which in turn may improve privacy protection. As such, we propose a data synthesis method entitled fidelity-agnostic synthetic data. The method first extracts features relevant to the dataset's intended use using a neural net and then generates synthetic versions of the extracted features, after which they are decoded to mimic the real dataset. We show that our synthetic data improve performance in prediction tasks while retaining privacy protection compared to other state-of-the-art methods.
Related Concept Videos
Distribution Reliability and Automation
Censoring Survival Data
Synthetic Biology
Golden rice
Golden rice is a genetically modified...
Propagation of Uncertainty from Random Error
Improving Translational Accuracy
Improving Translational Accuracy