Related Experiment Video
Updated: Apr 25, 2026

Assessing Early Stage Open-Angle Glaucoma in Patients by Isolated-Check Visual Evoked Potential
Published on: May 25, 2020
Comprehensive Evaluation of ChatGPT's Diagnostic Accuracy on Image-based Ophthalmic Case Interpretations
Cheng Jiao1, Iden Amiri1, Alice Yang Zhang1
1Department of Ophthalmology, University of North Carolina at Chapel Hill, Chapel Hill, North Carolina.
Objective:
To evaluate the diagnostic and treatment accuracy of Chat Generative Pre-trained Transformer (GPT-4.o) in ophthalmology, comparing performance when provided with full clinical context versus image-only inputs.
Design:
Cross-sectional diagnostic accuracy study.
Subjects:
A total of 261 ophthalmic cases from the University of Iowa's publicly available EyeRounds repository, spanning 6 subspecialties (pediatrics, retina, neuro-ophthalmology, glaucoma, oculoplastics, and cornea/cataract).
Methods:
Each case was analyzed under 2 conditions: (1) full-context input (clinical narrative and expert-generated image descriptions) and (2) image-only input (raw clinical images without text). Chat Generative Pre-trained Transformer was prompted to identify key signs/symptoms, differential diagnoses, final diagnosis, and treatment recommendations. Outputs were compared to expert-authored case solutions and evaluated for diagnostic accuracy and component match rates. Logistic regression was used to determine which output features predicted correct diagnosis.
Main Outcome Measures:
Primary outcome: diagnostic accuracy (match to expert diagnosis). Secondary outcomes: percentage match of signs/symptoms, differential diagnoses, and treatment recommendations.
Results:
Diagnostic accuracy was significantly higher with full-context input (80.1%) than image-only input (54.7%) (χ2 = 48.00, P < 0.001). Match rates for full versus image conditions were: signs/symptoms (56.5% vs. 42.6%), differential diagnoses (45.1% vs. 35.7%), treatment (64.8% vs. 49.6%), and overall match (53.8% vs. 41.5%) (all P < 0.001). In full-context cases, only treatment match significantly predicted diagnostic accuracy (odds ratio = 1.02, P < 0.001). In image-only cases, both signs/symptoms and treatment match predicted accuracy (odds ratio = 1.037 each, P < 0.001). Subspecialty analysis showed highest full-context accuracy in pediatrics (64.7%) and lowest image-only accuracy in glaucoma (21.7%) and neuro-ophthalmology (23.7%).
Conclusions:
Chat Generative Pre-trained Transformer demonstrated high diagnostic accuracy when provided with full clinical context but significantly reduced performance with image-only input. These findings highlight the limitations of multimodal artificial intelligence (AI) in interpreting raw ophthalmic images and underscore the importance of integrating structured clinical data to optimize AI-driven diagnostic support.
Financial Disclosures:
The author has no/the authors have no proprietary or commercial interest in any materials discussed in this article.

