Related Experiment Video
Updated: Jul 14, 2026

Quantitative Fundus Autofluorescence for the Evaluation of Retinal Diseases
Published on: March 11, 2016
A clinician aligned vision language framework for stepwise interpretation in fundus fluorescein angiography
Zichang Su1, Xiaocong Liu2, Bingtao Guan2
1Zhejiang University, Eye Center of Second Affiliated Hospital, School of Medicine, China. Zhejiang Provincial Key Laboratory of Ophthalmology. Zhejiang Provincial Clinical Research Center for Eye Diseases. Zhejiang Provincial Engineering Institute on Eye Diseases, Zhejiang University, Hangzhou, China.
None:
Fluorescein fundus angiography (FFA) is essential for diagnosing retinal vascular diseases, yet its interpretation is expertise-intensive. Here, we present Clin-FFA-VLM, a multimodal vision-language framework that mirrors retina specialists' cognitive workflow by decomposing FFA interpretation into three stages: lesion-aware visual perception, clinical report generation, and diagnostic decision support. Trained and tested on a multi-center dataset of 13,178 FFA images with expert label, 21,717 FFA images with 1790 clinical reports and diagnosis across 7 retinal diseases, Clin-FFA-VLM achieves an F1 of 0.834 for lesion detection, an entity-level F1 of 0.73 for report generation, and a diagnostic F1 of 0.77 by jointly reasoning over images and self-generated reports. External validation across two independent hospitals confirmed its generalizability (F1 of 0.78 and 0.70). In a prospective reader study with 200 FFA cases, Clin-FFA-VLM significantly improved diagnostic accuracy for medical students and residents (p < 0.05), bridging the gap between automated systems and clinical practice.
