Related Experiment Video
Updated: Aug 20, 2026

Functional Magnetic Resonance Imaging (fMRI) of the Visual Cortex with Wide-View Retinotopic Stimulation
Published on: December 8, 2023
Compact vision models match domain-specific foundation models for several retinal imaging classification tasks: A
Dávid Isztl1,2, Tahm Spitznagel1,2, Gábor Márk Somfai1,2
1Stadtspital Zurich, Department of Ophthalmology, Zurich, Switzerland.
None:
Large domain-specific foundation models have been widely adopted for retinal image analysis, yet systematic evidence for their advantage over compact general-purpose architectures remains scarce. We benchmarked nine model configurations spanning 22.8M to 303M parameters (vision transformers, hierarchical Swin Transformers, ConvNeXt, and the domain-specific RETFound models) across four tasks: OCT 8-class disease classification, and three fundus photography tasks (DME severity, glaucoma detection, and DR severity grading). All models were evaluated under identical training conditions, with both pretrained (on natural-domain image datasets) and from-scratch initializations compared using Mann-Whitney U tests. Pretraining improved accuracy by 5.18-18.41 percentage points across all tasks (p < 0.05 throughout), with larger benefits for CFP modalities and harder tasks. Compact hierarchical models (27-29M parameters) matched or exceeded larger architectures on three of four tasks. For instance, the SwinV2-tiny architecture ranked first on OCT, DME, and GL classification. The domain-specific RETFound model (303M) achieved the highest accuracy only on the most challenging task (DR severity grading, where the most severe class is underrepresented at 8% of images), where it outperformed the best compact model by 1.54 percentage points. These results indicate that compact general-purpose models may be sufficient for most retinal classification benchmarks, and that domain-specific foundation models may add higher value mainly for severity grading tasks with skewed class distributions.
