Related Experiment Video
Updated: May 14, 2026

09:17
Using Retinal Imaging to Study Dementia
Published on: November 6, 2017
Diagnostic Agreement Between a General-Purpose AI Model and Retinal Specialists in Color Fundus Photography-A Pilot
Sara Vaz-Pereira1,2, Laura Vilaverde3, André Ferreira3,4,5
1Department of Ophthalmology, Faculdade de Medicina, Universidade de Lisboa, 1649-004 Lisbon, Portugal.
Journal of Clinical Medicine
|May 13, 2026
Summary
This pilot study found that a general artificial intelligence (AI) model showed low agreement with retinal specialists when interpreting color fundus photographs. AI
Area of Science:
- Ophthalmology
- Medical Artificial Intelligence
- Computational Imaging
Background:
- Artificial intelligence (AI) shows promise in disease-specific retinal screening.
- Reliability of general-purpose AI in diverse clinical settings is uncertain.
- This study evaluates a multimodal AI model against human specialists for retinal image interpretation.
Purpose of the Study:
- To compare the diagnostic agreement of a general-purpose AI model with retinal specialists.
- To assess the reliability of AI in interpreting color fundus photographs (CFPs) in a clinical context.
- To evaluate AI performance in open-ended retinal image assessment.
Main Methods:
- Pilot retrospective cross-sectional study.
- 66 color fundus photographs (CFPs) evaluated by a masked retinal specialist and Google Gemini 2.5 Flash AI.
- Comparison against an unblinded treating specialist with full clinical information.
- Agreement assessed using weighted percent agreement and Gwet's AC2.
Main Results:
- Substantial agreement between two human specialists (AC2 = 0.67).
- Low agreement between the AI model and the reference specialist (AC2 = -0.58).
- Limited reliability in direct comparison between masked specialist and AI (AC2 = -0.38).
Conclusions:
- The evaluated AI model showed limited agreement with a context-informed specialist.
- Findings suggest cautious interpretation of consumer-facing AI for retinal image assessment.
- Further validation in larger, multicenter studies is warranted.