Related Experiment Videos
Hidden in Plain Sight: A Robust Generative Model for Secure Medical Data Sharing
Abstract:
The increasing use of AI for disease diagnosis necessitates sharing medical datasets across healthcare institutions, but this raises significant privacy concerns. Traditional methods like data de-identification, generative adversarial networks, and differential privacy offer privacy protection but involve risks and trade-offs in data utility, highlighting the need for more secure and efficient solutions. To tackle these problems, we introduce MedShare, a new dataset condensation approach that converts large datasets into a generative model. The process involves using an attention-based generative adversarial network (GAN), followed by leveraging a vision transformer to capture a refined set of dataset features. This method is capable of generating synthetic medical images that reflect the original data and facilitate secure sharing to enhance downstream task performance across hospitals. Evaluations on seven public datasets with diverse data modalities demonstrate that MedShare outperforms baseline methods by approximately 5-30%. A synthetic dataset generated by MedShare, representing only 5% of the original data, achieves 96% of the Area Under the Curve (AUC) performance of the original dataset. Furthermore, MedShare strengthens privacy by minimizing data leakage and offering robust defenses against privacy attacks on both data and models, showcasing its significant potential for safeguarding patient privacy. These advancements affirm its effectiveness in promoting secure sharing of medical data and models, fostering innovation in AI-generated content for healthcare.