Related Experiment Video
Updated: May 7, 2026

07:22
Quantitative Fundus Autofluorescence for the Evaluation of Retinal Diseases
Published on: March 11, 2016
11.9K
Can Multimodal Large Language Models Diagnose Diabetic Retinopathy from Fundus Photos? A Quantitative Evaluation
Jesse A Most1,2, Evan H Walker3, Nehal N Mehta1,3
1Jacobs Retina Center, Shiley Eye Institute, University of California San Diego, La Jolla, California.
Ophthalmology Science
|October 1, 2025
Summary
Four multimodal large language models (MLLMs) showed variable accuracy in detecting diabetic retinopathy (DR). While some models performed well in specific tasks, their overall diagnostic performance requires further improvement for clinical use.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Imaging
Background:
- Diabetic retinopathy (DR) is a leading cause of vision loss.
- Early detection and grading of DR are crucial for effective management.
- Advancements in AI, particularly multimodal large language models (MLLMs), offer potential for automated image analysis.
Purpose of the Study:
- To evaluate the diagnostic accuracy of four MLLMs in detecting and grading diabetic retinopathy (DR).
- To assess the performance of MLLMs using their image analysis capabilities on fundus images.
Main Methods:
- A retrospective study analyzed 309 ultra-widefield fundus images from 188 patients with diabetes or prediabetes.
- Four MLLMs (ChatGPT-4o, Claude 3.5 Sonnet, Google Gemini 1.5 Pro, Perplexity Llama 3.1 Sonar/Default) were tested with distinct prompts for DR diagnosis and severity grading.
- Ground truth was established by three retina specialists using the ETDRS classification system.
Main Results:
- For multi-choice DR identification, Claude and ChatGPT demonstrated higher accuracy and sensitivity compared to other models.
- In binary DR classification, ChatGPT and Perplexity achieved the highest accuracy, with all models showing high specificity.
- For DR severity grading, no significant differences in accuracy were found between models, and all exhibited low sensitivity.
Conclusions:
- MLLMs show promise for assisting in retinal image analysis for DR detection and grading.
- Current diagnostic performance of MLLMs does not meet clinical standards for safe implementation.
- Further model training and error optimization are necessary to enhance the clinical utility of MLLMs in DR management.

