Related Experiment Video
Updated: Apr 18, 2026

Assessing Early Stage Open-Angle Glaucoma in Patients by Isolated-Check Visual Evoked Potential
Published on: May 25, 2020
Comparative assessment of diagnostic agreement between artificial intelligence and general practitioners in diabetic
María Alcalá-Millán1, Juan Carlos Romero-Vigara2, Eva Trillo-Calvo3
1Optometrist Primary Care, Aragon Institute for Health Research (IIS Aragón), Primary Care Centre Sagasta-Ruiseñores, Zaragoza, Spain.
Aims:
To assess the diagnostic agreement between an artificial intelligence (AI) system and general practitioners (GPs) interpreting fundus photographs for diabetic retinopathy (DR) screening, using the ophthalmologists' assessment as the reference standard.
Methods:
We performed a cross-sectional study of 500 primary care patients with type 2 diabetes (T2DM). Each underwent two 45° non-mydriatic fundus photos per eye. Images were independently evaluated by three trained GPs, a deep learning-based AI system (EyeArt v3.0.0 (Eyenuk Inc.)), and an ophthalmologist (reference standard). Diagnostic agreement was measured with Cohen's kappa and Cramer's V, and sensitivity, specificity, predictive values, likelihood ratios, and overall accuracy were calculated.
Results:
Mean age was 64 years, and 59% were men. DR prevalence was 11%. AI showed near-perfect agreement with ophthalmology (κ = 0.91) and outperformed GPs, with 98.2% sensitivity, 98.9% specificity, and 98.8% accuracy. GPs showed substantial agreement (κ = 0.84), with lower sensitivity (86.4%) but similarly high specificity (98.1%). In multinomial analysis, AI achieved 87.6% sensitivity for mild DR and 100% for moderate-to-severe DR, missing no clinically relevant cases.
Conclusions:
In this cohort restricted to gradable images, AI demonstrated higher sensitivity, specificity, and diagnostic agreement than GPs when compared with a single ophthalmologist reference standard. Supervised use in primary care could strengthen population-based screening, but large-scale adoption will require multicentre studies, formal cost-effectiveness analyses, and validation in unselected screening populations.

