Related Experiment Video
Updated: Aug 6, 2026

Assessing Early Stage Open-Angle Glaucoma in Patients by Isolated-Check Visual Evoked Potential
Published on: May 25, 2020
Comparative performance of chatgpt and gemini in diagnostic classification and clinical reasoning for open-angle
Zhewen Zhang1, Siyu Lu2, Zhenqiang Xu3
1Department of Ophthalmology, Linping District Hospital of Traditional Chinese Medicine, No. 101, Yuncheng Street, Hangzhou, 311100, China. 17757726376@163.com.
Purpose:
To evaluate differences in performance between two large language models (LLMs), GPT-5.3 and Gemini 2.5 Pro, in diagnostic classification and clinical reasoning for primary open-angle glaucoma (POAG).
Methods:
Forty-eight guideline-based standardized cases were constructed. Using a unified prompt, we fed the cases into both models and obtained outputs on diagnosis, classification, and reasoning. A consensus of three glaucoma specialists served as the reference standard. We compared the diagnostic accuracy and classification consistency (Cohen's κ) of the two models. Clinical reasoning ability was scored using a Likert scale based on logical coherence, evidence utilization, and conclusion consistency. We further analyzed error patterns and potential safety issues.
Results:
Overall diagnostic accuracy was 85.4% for GPT-5.3 and 75.0% for Gemini (P = 0.306). Both models performed well on typical cases but showed reduced accuracy on borderline cases, including ocular hypertension, suspected glaucoma, and early-stage POAG. For classification consistency, κ values were 0.675 for GPT-5.3 and 0.628 for Gemini. GPT-5.3 scored higher than Gemini in overall clinical reasoning (4.4 ± 0.6 vs. 3.9 ± 0.7, P = 0.011). Error pattern analysis indicated that Gemini was more prone to overdiagnosis and reasoning inconsistency, whereas GPT-5.3 was relatively conservative. Both models had low rates of unsafe outputs, though Gemini showed a slightly higher proportion.
Conclusion:
ChatGPT and Gemini both demonstrate certain capabilities in diagnosing POAG, but their stability on borderline cases remains limited. Comparatively, GPT-5.3 shows higher consistency and more stable reasoning patterns. The application of LLMs in ophthalmic diagnostic support still requires cautious evaluation.
Related Concept Videos
Open Angle Glaucoma: Treatment
Drugs such as carbonic anhydrase inhibitors, α2- and...
Glaucoma: Overview
Angle Closure Glaucoma: Treatment