Related Experiment Video
Updated: Jun 17, 2026

Comparison of Agreement and Accuracy using Binocular Wavefront Optometer with Autorefractor and Phoropter
Published on: September 16, 2025
Comparative Accuracy and Safety of 4 Large Language Models on Cornea and External Disease Multiple-Choice Questions
Bora Deniz Argon1, Şule Vildan Durmuş
1Department of Ophthalmology, University of Health Sciences, Prof. Dr. Cemil Taşcıoğlu City Hospital, Şişli, Istanbul, Turkey.
Purpose:
To compare the accuracy and safety of 4 large language models on cornea and external disease multiple-choice questions (MCQs).
Methods:
DeepSeek-V3.2, GPT-5.2 (via ChatGPT), Gemini 3 Pro, and Claude Opus 4.5 were tested on American Academy of Ophthalmology (AAO) Ophthalmic News and Education (ONE) Network Cornea/External MCQs (114 text-only) and AAO Basic and Clinical Science Course Cornea/External study questions (38 text-only). Each item was queried once per model using a standardized prompt for the primary analysis. In secondary analyses, each item was requeried 5 times per model in independent sessions. Accuracy was compared using Cochran Q and pairwise exact McNemar tests with Holm adjustment. Potentially harmful wrong answers (HarmfulWrong) were independently coded using a prespecified rubric.
Results:
In the AAO ONE dataset, accuracy was 65.8% for DeepSeek-V3.2, 93.0% for GPT-5.2, 93.9% for Gemini 3 Pro, and 92.1% for Claude Opus 4.5 (Cochran Q P < 0.001). DeepSeek-V3.2 was significantly less accurate than all other models; the other 3 did not differ significantly. In the Basic and Clinical Science Course dataset, accuracy was 57.9%, 76.3%, 94.7%, and 84.2%, respectively (Cochran Q P < 0.001); DeepSeek-V3.2 was significantly less accurate than Gemini 3 Pro and Claude Opus 4.5. Five-run retesting showed similar ranking but imperfect stability, particularly for DeepSeek-V3.2. HarmfulWrong responses were infrequent, and 2 cornea specialists achieved 98.0% and 97.4% overall accuracy without HarmfulWrong responses.
Conclusions:
GPT-5.2, Gemini 3 Pro, and Claude Opus 4.5 achieved similarly high accuracy on text-only cornea/external disease MCQs, whereas DeepSeek-V3.2 underperformed. Potentially harmful errors were uncommon but support continued clinician oversight.
