Related Experiment Video
Updated: Jun 16, 2025

04:23
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
1.8K
Comparative Analysis of Vision Transformers and Conventional Convolutional Neural Networks in Detecting Referable
Jocelyn Hui Lin Goh1, Elroy Ang2, Sahana Srinivasan1
1Singapore Eye Research Institute, Singapore National Eye Center, Singapore, Singapore.
Ophthalmology Science
|August 21, 2024
Summary
Vision transformers (ViTs) significantly outperform convolutional neural networks (CNNs) in detecting referable diabetic retinopathy (DR). The SWIN transformer model demonstrated superior performance in identifying DR from retinal images.
Area of Science:
- Ophthalmology
- Medical Imaging
- Artificial Intelligence
Background:
- Diabetic retinopathy (DR) detection from retinal images is crucial for preventing vision loss.
- Convolutional neural networks (CNNs) have dominated image classification tasks, including DR detection.
- Vision transformers (ViTs) show promise but their efficacy in referable DR detection remains underexplored.
Purpose of the Study:
- To compare the performance of Vision transformers (ViTs) and Convolutional Neural Networks (CNNs) for referable diabetic retinopathy (DR) detection.
- To evaluate the diagnostic accuracy of different ViT and CNN models using retinal photographs.
Main Methods:
- A retrospective study involving 48,269 retinal images from Kaggle, Messidor-1, and SEED datasets.
- Development and comparison of 5 CNN models and 4 ViT models for referable DR detection.
- Performance evaluation using Area Under the Operating Characteristics Curve (AUC), specificity, and sensitivity on internal and external test sets.
Main Results:
- The SWIN transformer (a ViT model) achieved the highest AUC (95.7% internal, 97.3% SEED, 96.3% Messidor-1), significantly outperforming all CNN models.
- At 80% specificity, the SWIN transformer demonstrated superior sensitivity (94.4%) compared to CNNs (76.3%-83.8%).
- These performance advantages were consistent across internal and external validation datasets.
Conclusions:
- Vision transformers (ViTs) offer superior performance compared to CNNs for detecting referable diabetic retinopathy from retinal images.
- ViT models, particularly the SWIN transformer, hold significant potential for enhancing deep learning-based DR detection systems.
- These findings suggest ViTs can optimize automated analysis of retinal photographs for improved DR screening and management.

