Related Experiment Video
Updated: Jan 8, 2026

Assessing Early Stage Open-Angle Glaucoma in Patients by Isolated-Check Visual Evoked Potential
Published on: May 25, 2020
Comparative performance of ChatGPT o3-mini-high and DeepSeek-R1 in ophthalmology: An evaluation of diagnostic
G Karataş1, F Kırık2, M E Karataş3
1Department of Ophthalmology, Prof. Dr. Cemil Taşcıoğlu City Hospital, Istanbul, Turkey.
Purpose:
To compare the diagnostic reasoning and case-based problem-solving abilities of ChatGPT o3-mini-high and DeepSeek-R1 in ophthalmological cases with text-based questions.
Methods:
Fifty-five consecutive text-based case-solving questions from nine ophthalmology subspecialties were posed to two reasoning-capable LLMs, ChatGPT o3-mini-high and DeepSeek-R1. For each case, the multi-component diagnostic reasoning approach described by Elstein was applied. Overall diagnostic accuracy, diagnostic agreement between the models, reasoning competence, subspecialty-specific performance, and the tendency to request additional prompts were recorded. Two expert ophthalmologists then independently rated each model's diagnostic reasoning ability for all questions, using the Global Quality Score, a five-point scale (1=poor; 5=excellent).
Results:
ChatGPT o3-mini-high correctly answered 80% of the questions, whereas DeepSeek-R1 achieved a correct response rate of 54.5% (P<0.001), and Cohen's kappa coefficient was 0.462. ChatGPT o3-mini-high tended to request additional prompts for responses to fewer questions (2 vs. 12; P: 0.013). For both LLMs, the highest accuracy was observed in the retina/vitreous-related cases, while the lowest accuracy was noted in glaucoma-related cases. When Elstein's medical reasoning components were evaluated with the GQS, ChatGPT o3-mini-high achieved a median score of 4.5 (IQR 2.5-5.0), whereas DeepSeek-R1 achieved 2.5 (IQR 1.0-4.5) (P<0.001). The weighted kappa was 0.407, indicating moderate agreement between the two models.
Conclusion:
This study provides evidence that ChatGPT o3-mini-high demonstrates superior diagnostic accuracy and reasoning capabilities in the analysis of ophthalmologic cases compared to DeepSeek-R1.

