Related Experiment Videos
Privacy-aware diabetic retinopathy grading and visual lesion-focused interpretability through mixture-of-experts
Md Tanjum An Tashrif1, Dipanjali Kundu1, Mst Moriom Akter Bithee1
1Department of CSE, National Institute of Textile Engineering and Research (NITER), Constituent Institute of the University of Dhaka, Savar, Dhaka, 1350, Bangladesh.
Scientific Reports
|June 23, 2026
Summary
A new federated Mixture-of-Experts (FL-MoE) framework enhances automated Diabetic Retinopathy (DR) screening. This privacy-preserving system improves model performance and stability across diverse datasets, offering a practical solution for clinical settings.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Computer Vision
Background:
- Diabetic Retinopathy (DR) is a leading preventable cause of vision loss.
- Automated screening systems are needed to analyze medical data across institutions without centralizing sensitive information.
- Federated Learning (FL) enables collaborative model training while maintaining data privacy, but faces challenges with non-IID data, communication overhead, and convergence stability in medical imaging.
Purpose of the Study:
- To propose and evaluate a federated Mixture-of-Experts (FL-MoE) framework for Diabetic Retinopathy classification.
- To combine interpretable deep learning with expert specialization to address FL limitations in medical imaging.
- To assess the performance of various backbone architectures (CNN, CNN-LSTM, ViT) within the FL-MoE framework using public DR datasets.
Main Methods:
- Developed a federated Mixture-of-Experts (FL-MoE) framework for DR classification.
- Evaluated multiple deep learning backbone architectures (CNN, CNN-LSTM, Vision Transformer) within the FL-MoE setup.
- Utilized the EyePACS and APTOS-2019 retinal fundus datasets for training and validation.
- Employed Grad-CAM for explainability analysis and Intersection-over-Union (IoU) for localization quality assessment.
Main Results:
- The FL-MoE framework demonstrated improved performance under heterogeneous data distributions across different backbone architectures.
- The CNN-LSTM backbone achieved 76.2% accuracy and 89.5% AUC on the EyePACS dataset, with significantly reduced communication costs compared to transformer models.
- CNN-LSTM exhibited more stable convergence and greater robustness to client-level data heterogeneity.
- Explainability analysis showed attention maps highlighting relevant retinal regions, though localization accuracy (mean IoU < 0.03) was coarse.
Conclusions:
- The proposed FL-MoE framework with a CNN-LSTM backbone provides an effective and scalable solution for privacy-aware Diabetic Retinopathy screening.
- This approach outperforms standard federated baselines and personalized FL methods under heterogeneous data conditions.
- The framework offers a practical method for federated clinical environments, enhancing diagnostic capabilities while preserving patient data privacy.