Addressing the Generalizability of AI in Radiology Using a Novel Data Augmentation Framework with Synthetic Patient
Gianluca Brugnara1, Chandrakanth Jayachandran Preetha1, Katerina Deike1
1From the Department of Neuroradiology (G.B., C.J.P., M.F.D., M.A.M., M.B., H.M., A. Rastogi, P.V.), Division for Computational Neuroimaging (G.B., C.J.P., M.F.D., M.A.M., H.M., A. Rastogi, P.V.), and Department of Neurology (B.W., R.D., W.W.), Heidelberg University Hospital, Im Neuenheimer Feld 400, 69120 Heidelberg, Germany; Department of Neuroradiology (G.B., K.D., R.H., M.F.D., A. Radbruch, P.V.), Division for Computational Radiology and Clinical AI (G.B., M.F.D., A. Radbruch, P.V.), Bonn University Hospital, Bonn, Germany; German Center for Neurodegenerative Diseases (DZNE), Bonn, Germany (K.D., A. Radbruch); Division of Medical Image Computing, German Cancer Research Center (DKFZ), Heidelberg, Germany (G.B., P.V.); and Institute for Applied Mathematics, University of Bonn, Bonn, Germany (T.P.).
Generative adversarial networks (GANs) create synthetic MRI data to improve artificial intelligence (AI) model performance on new patient datasets. This synthetic data augmentation enhances AI generalizability for multiple sclerosis lesion detection.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Machine Learning
Background:
- Artificial intelligence (AI) models often experience performance degradation when applied to external datasets due to domain shift.
- Generalizability of AI models in medical imaging is crucial for reliable clinical application, particularly in longitudinal studies.
- Multiple Sclerosis (MS) lesion detection in MRI requires robust AI models that can perform consistently across different data acquisition parameters.
Purpose of the Study:
- To evaluate a novel generative adversarial network (GAN)-based data augmentation framework for creating synthetic patient MRI data.
- To improve the generalizability and robustness of AI models for detecting new MS lesions in MRI during longitudinal follow-up.
- To assess the impact of synthetic data augmentation on AI model performance when tested on external datasets from a different institution.
Main Methods:
- An attention-based AI network was developed using an internal dataset of 669 MS patients (3083 examinations).
- The network was trained both with and without the inclusion of GAN-generated synthetic MRI data.
- External validation was performed on 134 MS patients from a different institution, using data from varied scanners and protocols.
Main Results:
- Models trained with GAN-based synthetic data augmentation showed a significant performance increase on external data (AUC 93.3% vs 83.6%, P = .03).
- Synthetic data augmentation resulted in performance comparable to the internal test set (AUC 93.3% vs 95.0%, P = .53).
- Models without synthetic data augmentation exhibited a performance drop on external data (AUC 83.6% vs 93.8% on internal data, P = .03).
Conclusions:
- Data augmentation using GAN-generated synthetic patient data substantially enhances AI model performance on unseen MRI data.
- This approach effectively mitigates performance drops caused by domain shift in external datasets.
- Synthetic data augmentation holds potential for improving AI robustness in medical imaging across various clinical conditions and tasks.


