Related Experiment Video
Updated: Apr 10, 2026

Author Spotlight: Ex Vivo OCT-Based Multimodal Imaging of Human Donor Eyes for Research into Age-Related Macular Degeneration
Published on: May 26, 2023
MaxGRNet: A multi-axis vision transformer with improved generalization for eye disease classification using
Md Mehedi Hasan Santo1, Fuyad Hasan Bhoyan2, Fuad Ibne Jashim Farhad1
1Department of Information Technology, Central Queensland University, Melbourne, Victoria, Australia.
Abstract:
Eye diseases, including diabetic retinopathy (DR), glaucoma, and cataracts, represent a major global health concern and can lead to severe visual impairment or blindness if not identified in a timely manner. This study proposes a novel eye disease classification framework based on a multi-axis vision transformer (MaxViT) applied to color fundus images with Explainable Artificial Intelligence (XAI) techniques to enhance model transparency. The proposed architecture integrates transformer-based attention mechanisms with Global Response Normalization (GRN)-based multi-layer perceptron (MLP) layers to capture complex spatial and contextual relationships within fundus images effectively. The model was evaluated on a publicly available eye disease classification dataset using a five-fold cross-validation strategy to assess its robustness and generalization. The experimental results show that the proposed approach consistently outperforms conventional Convolutional Neural Networks (CNNs) and Vision Transformer (ViT) variants, including ResNet50, Swin-T, MaxViT-T, and ViT-B16. The model achieved a macro-averaged test accuracy, precision, and recall values of 96.75%, 96.70%, and 96.80%, respectively, with paired statistical t-tests confirming that these improvements were statistically significant. Rigorous preprocessing techniques were employed to improve data consistency, and XAI-based visual explanations provided insights into the model's decision-making process, supporting interpretability in ophthalmic image analysis. Overall, the proposed MaxViT-based framework is robust and computationally feasible for research-oriented evaluation approaches for automated fundus image classification, highlighting the potential of advanced transformer architectures for future decision-support and research-oriented ophthalmic applications.
More Related Videos
12:48In Vivo Dynamics of Retinal Microglial Activation During Neurodegeneration: Confocal Ophthalmoscopic Imaging and Cell Morphometry in Mouse Glaucoma
Published on: May 11, 2015
04:48Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022