Related Experiment Video
Updated: Aug 27, 2026

Using Retinal Imaging to Study Dementia
Published on: November 6, 2017
Diagnostic Accuracy of Multimodal Large Language Models in Retinal Fundus Photography
Lia Huo1, Astha Chandra2, Michael Balas3
1Institute of Medical Science, University of Toronto, Toronto, ON, Canada.
Purpose:
Artificial intelligence in ophthalmology has mostly used task-specific convolutional neural networks. The performance of general-purpose multimodal large language models (LLMs) for fundus image diagnosis remains largely unexplored. This study evaluated and compared the performance of four multimodal LLMs in diagnosing retinal diseases from color fundus photographs.
Methods:
A total of 800 images (100 per category across 8 conditions) were randomly sampled and evaluated by ChatGPT-4o, Claude 3.5 Sonnet, Gemini 1.5, and Grok-2. Each model received a standardized prompt with a single fundus photograph and was asked to provide a free-text diagnosis, confidence rating (1-10), and justification type, defined as the reasoning category behind its diagnosis. Primary outcome was overall diagnostic accuracy. Secondary outcomes included sensitivity, specificity, F1-scores, misclassification patterns, justification-accuracy association, confidence-accuracy correlation, Cohen's κ agreement across models, and mean response time per image.
Results:
ChatGPT achieved the highest diagnostic accuracy (75%; 603/800), far exceeding Claude (18%), Grok (18%), and Gemini (14%). Unlike other models, ChatGPT demonstrated a meaningful association between justification type and diagnostic correctness and showed a positive correlation between confidence and accuracy (r = 0.19; P < 0.001). Claude and Gemini exhibited weaker confidence-accuracy correlations, while Grok showed none. Agreement across models was consistently low (κ = 0.11-0.16 with ChatGPT; κ ≈ 0.05 among Claude, Gemini, and Grok).
Conclusions:
ChatGPT demonstrated the strongest accuracy, justification-accuracy association, and confidence-accuracy correlation compared with Claude, Gemini, and Grok. While promising, refinement, validation, and workflow integration are essential before clinical deployment.