Related Experiment Video
Updated: May 17, 2025

13:44
Detection of Architectural Distortion in Prior Mammograms via Analysis of Oriented Patterns
Published on: August 30, 2013
42.6K
How Foundational Is the Retina Foundation Model? Estimating RETFound's Label Efficiency on Binary Classification of
David Kuo1, Qitong Gao2, Dev Patel2
1Department of Ophthalmology, Duke University, Durham, North Carolina.
Ophthalmology Science
|March 31, 2025
Summary
The RETFound model, pretrained on extensive ophthalmology data, significantly outperforms other models like ResNet-50 and ViT-L in identifying normal versus abnormal OCT B-scans. This demonstrates the value of specialized pretraining for medical imaging tasks, even with limited labeled data.
Area of Science:
- Ophthalmology
- Medical Imaging
- Machine Learning
Background:
- Machine learning progress is often limited by the scarcity of labeled medical data due to privacy and cost constraints.
- Self-supervised learning, particularly pretraining on large unlabeled datasets, shows promise for improving model performance on downstream tasks.
- The RETFound model, a large vision transformer pretrained on 1.6 million retinal images, represents a significant advancement in ophthalmology AI.
Purpose of the Study:
- To evaluate the label efficiency of the RETFound model for identifying normal versus abnormal Optical Coherence Tomography (OCT) B-scans.
- To compare the performance of RETFound against established models (ResNet-50, ViT-L) in a primary care diabetic retinopathy screening context.
Main Methods:
- 1150 OCT B-scans from 647 patients were randomly allocated into training, validation, and testing sets.
- Three models (ResNet-50, ViT-L, RETFound) were fine-tuned on varying dataset sizes (915 to 50 OCT B-scans).
- Model performance was assessed using accuracy, AUROC, AUPRC, F1 score, precision, and recall.
Main Results:
- RETFound consistently outperformed ResNet-50 and ViT-L across all metrics and dataset sizes.
- ResNet-50 and ViT-L showed comparable performance on larger datasets but degraded significantly with smaller subsets.
- The study highlights RETFound's superior ability to learn from limited labeled data, emphasizing the benefit of retina-specific pretraining.
Conclusions:
- The findings validate the effectiveness of the RETFound model's specialized pretraining for ophthalmology tasks.
- RETFound demonstrates strong label efficiency, making it suitable for applications with limited labeled OCT B-scan data.
- Further research is recommended to optimize fine-tuning strategies for RETFound in various downstream applications.

