Related Experiment Videos
A generalizability analysis framework for bias resilience in convolution neural network and transformer models for
Majida Kazmi1,2, Bisma Imran2, Saad A Qazi1,2
1Faculty of Electrical and Computer Engineering, NED University of Engineering & Technology, Karachi, Pakistan.
None:
Automated diabetic retinopathy (DR) screening has seen extensive research progress, yet its clinical adoption remains limited, largely due to insufficient attention to model generalizability across diverse populations and imaging conditions. Existing studies often overlook the role of dataset biases, arising from demographic, acquisition, and preprocessing variations that undermine robustness in real-world settings. To address this critical gap, we propose a comprehensive generalizability analysis framework that systematically categorizes dataset biases, evaluates their impact on performance, and introduces robust inter-dataset comparison metrics. The framework is metric-agnostic and extensible to multi-class severity grading, with stage-wise BAG formulations introduced to support more detailed clinical evaluation. Using one primary (EyePACS) and nine secondary datasets, we assessed the generalizability of two distinct architectures: a convolutional neural network (MobileNetV2) and a transformer-based model (CvT). Performance was evaluated through intra-group measures (accuracy, sensitivity, specificity, and AUC) and a newly proposed Bias-Adjusted Generalization (BAG) index designed to quantify resilience to bias-induced domain shifts. Results demonstrate that CvT consistently outperformed MobileNetV2, achieving an average accuracy of 87% and a BAG index of 0.94, compared to MobileNetV2's 78% accuracy and BAG index of 0.91. Gradient-based explainability visualizations further confirm that both models attend to clinically relevant retinal regions, with architectural differences in activation patterns consistent with their respective generalizability profiles. These findings highlight the inherent bias resilience of transformer architectures, while highlighting their computational demands as a barrier to deployment in low-resource environments. Importantly, the study suggests that integrating transformer-inspired mechanisms into lightweight CNNs could yield clinically scalable models with both efficiency and robustness. Our framework offers a standardized approach for bias-aware evaluation, providing actionable insights for developing equitable, generalizable, and resource-adaptable AI solutions for DR screening.