Related Experiment Videos
NoisyFlow: differentially private optimal transport using neural networks for secure biomedical data sharing across
Yunyang Li1, Nikhil Khandekar1, Skylar Wang1
1Department of Computer Science, Yale University, New Haven, CT 06511, United States.
Bioinformatics (Oxford, England)
|July 7, 2026
Summary
NoisyFlow enables privacy-preserving multi-institutional biomedical analysis by harmonizing data across sites. This differentially private framework reduces distribution shift, improving model performance without sharing sensitive patient records.
Area of Science:
- Biomedical data science
- Computational privacy
- Machine learning for healthcare
Background:
- Sharing sensitive patient data across institutions is crucial for improving biomedical models but faces privacy constraints.
- Data distribution shifts (e.g., batch effects) between institutions hinder multi-institutional analyses.
- Existing privacy-preserving methods do not adequately address cross-site distribution mismatch.
Purpose of the Study:
- To develop a privacy-preserving framework for harmonizing multi-institutional biomedical data under distribution shift.
- To enable reliable cross-institutional analyses without compromising patient privacy.
- To address the challenge of integrating data from diverse sources while maintaining data utility.
Main Methods:
- NoisyFlow is a three-stage differentially private framework.
- Stage I: Each site learns a differentially private flow-based generator of its local distribution.
- Stage II: A neural optimal transport map aligns local distributions to a shared reference.
- Stage III: A central server composes models to generate reference-aligned pseudo-data.
Main Results:
- NoisyFlow effectively reduces distribution shift across different biomedical settings.
- The framework preserves downstream analytical utility under formal differential privacy guarantees.
- Demonstrated success in single-cell genomics, histopathology, neurogenomics, and wearable sensing data.
Conclusions:
- NoisyFlow provides a robust solution for privacy-preserving cross-institutional data harmonization.
- The method facilitates reliable multi-institutional analyses by addressing both privacy and distribution shift.
- Enables the development of more generalizable and accurate biomedical models through pooled, harmonized data.