Related Experiment Videos
SSLVNet: an explainable multi-modal hybrid framework integrating self-supervised learning and vision transformers for
1School of Computer Science Engineering and Information Systems, Vellore Institute of Technology, Vellore, Tamil Nadu, India.
Introduction:
Retinal Diseases are the main reason for vision loss and blindness worldwide. Initial identification of retinal abnormalities is essential for avoiding severe retinal-related complications and improving patient results. Traditional retinal disease diagnosis mainly depends on manual checks of retinal fundus images, which are slow and highly dependent on medical expertise. Deep learning methods are used in retinal disease classification. Existing methods are primarily based on labelled image data and often do not fully utilize complementary clinical information. The proposed hybrid multi-model framework classifies retinal diseases from the available retinal fundus images.
Methods:
SSLVNet combines a Self-Supervised Learning encoder, Lesion Attention, CNN, ViT, and a Simplified lesion Graph. Also, Grad- CAM provides visual elucidations of the model predictions. ODIR-5K dataset, which includes multiple ocular disease categories, is used.
Results:
The results are promising, with an Exact Accuracy of 84.44%, a Hamming accuracy of 96.92%, and a Macro F1-Score of 90.36%. An ablation study was done to understand the contribution of individual components in the proposed framework.
Discussion:
These results indicate that the proposed work can effectively support reliable and interpretable retinal disease classification in real-world healthcare applications.