Related Experiment Video
Updated: Apr 10, 2026

Author Spotlight: Ex Vivo OCT-Based Multimodal Imaging of Human Donor Eyes for Research into Age-Related Macular Degeneration
Published on: May 26, 2023
MaxGRNet: A multi-axis vision transformer with improved generalization for eye disease classification using
Md Mehedi Hasan Santo1, Fuyad Hasan Bhoyan2, Fuad Ibne Jashim Farhad1
1Department of Information Technology, Central Queensland University, Melbourne, Victoria, Australia.
A new multi-axis vision transformer (MaxViT) framework accurately classifies eye diseases from fundus images. This advanced AI model offers significant improvements over existing methods for early detection and diagnosis.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Computer Vision
Background:
- Eye diseases like diabetic retinopathy, glaucoma, and cataracts pose significant global health risks, potentially leading to blindness.
- Timely identification of these conditions is crucial for effective treatment and vision preservation.
- Current automated methods for analyzing ophthalmic images require further enhancement in accuracy and transparency.
Purpose of the Study:
- To introduce a novel eye disease classification framework utilizing a multi-axis vision transformer (MaxViT).
- To integrate Explainable Artificial Intelligence (XAI) techniques for enhanced model transparency and interpretability.
- To evaluate the framework's performance against established deep learning models for ophthalmic image analysis.
Main Methods:
- Development of a MaxViT architecture incorporating transformer attention and GRN-based MLP layers for complex feature extraction.
- Application of the framework to a public dataset of color fundus images for eye disease classification.
- Rigorous data preprocessing and a five-fold cross-validation strategy to ensure model robustness and generalizability.
- Evaluation using metrics such as accuracy, precision, and recall, with statistical significance testing.
Main Results:
- The proposed MaxViT framework achieved superior performance compared to CNNs (ResNet50) and ViT variants (Swin-T, MaxViT-T, ViT-B16).
- Macro-averaged test accuracy, precision, and recall reached 96.75%, 96.70%, and 96.80%, respectively, with statistically significant improvements.
- XAI techniques provided valuable insights into the model's decision-making process, enhancing interpretability.
Conclusions:
- The MaxViT-based framework demonstrates robust and computationally feasible performance for automated fundus image classification.
- This approach holds significant potential for advancing decision-support systems and research in ophthalmic applications.
- The integration of advanced transformer architectures offers a promising direction for future automated eye disease diagnosis.
More Related Videos
12:48In Vivo Dynamics of Retinal Microglial Activation During Neurodegeneration: Confocal Ophthalmoscopic Imaging and Cell Morphometry in Mouse Glaucoma
Published on: May 11, 2015
04:48Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022