Related Experiment Video
Updated: Sep 30, 2026

Using Retinal Imaging to Study Dementia
Published on: November 6, 2017
Resolution-Aware Cross-Modality Evaluation of CFP-Pretrained Retinal Vision-Language Models for Ultra-Widefield
Siyun Lee1, Baek Hwan Cho2,3,4,5, Jihyeon Son2
1Department of Ophthalmology, CHA University Bundang Medical Center, CHA University School of Medicine, Seongnam, Republic of Korea.
Abstract:
Retinal vision-language models (VLMs) pretrained on conventional color fundus photography (CFP) may be applied to other imaging systems without adequate modality-specific validation. We hypothesized that direct zero-shot transfer to ultra-widefield (UWF) photographs would produce class-specific failure and assessed whether linear probing showed broader five-grade diabetic retinopathy (DR) separation in a held-out set. In this retrospective single-center study, 300 de-identified UWF images from 120 eyes of 61 patients with diabetes were graded as no DR, mild, moderate, or severe nonproliferative DR, or proliferative DR. FLAIR and CLIP-DR were assessed at 512 × 512, 768 × 768, and 1024 × 1024 pixels. Both strategies were evaluated on the same 151-image test set from 24 patients; linear probes were fitted using a separate 149-image training set from 37 patients. Balanced accuracy, macro-F1, and macro-area under the receiver operating characteristic curve (macro-AUROC) were reported for both models. Class-specific and ordinal-error measures were also reported. CLIP-DR linear probing achieved macro-F1 of 0.365 at 1024 × 1024, compared with 0.254 for zero-shot inference. FLAIR zero-shot failed to identify severe nonproliferative DR at every resolution. FLAIR linear-probing macro-F1 peaked at 0.338 at 768 × 768 and decreased to 0.295 at 1024 × 1024; moderate- and severe-stage recall remained low. Direct CFP-to-UWF reuse therefore produced marked class collapse. Linear probing yielded higher macro-F1 estimates, but patient-cluster bootstrap intervals for all four primary contrasts included zero. Increasing resolution did not consistently improve performance. Modality-specific, class-resolved validation is required before retinal VLMs are reused across fundus imaging systems.
