Related Experiment Video
Updated: Jun 17, 2026

05:14
Comparison of Agreement and Accuracy using Binocular Wavefront Optometer with Autorefractor and Phoropter
Published on: September 16, 2025
Comparative Accuracy and Safety of 4 Large Language Models on Cornea and External Disease Multiple-Choice Questions
Bora Deniz Argon1, Şule Vildan Durmuş
1Department of Ophthalmology, University of Health Sciences, Prof. Dr. Cemil Taşcıoğlu City Hospital, Şişli, Istanbul, Turkey.
Cornea
|June 16, 2026
Summary
Four large language models were evaluated on ophthalmology questions. GPT-5.2, Gemini 3 Pro, and Claude Opus 4.5 demonstrated high accuracy, while DeepSeek-V3.2 performed significantly lower, indicating a need for continued clinician oversight.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Education
Background:
- Large language models (LLMs) show promise in various fields, including medicine.
- Evaluating LLM performance on specialized medical knowledge is crucial for safe integration.
- Cornea and external disease is a subspecialty within ophthalmology requiring precise knowledge.
Purpose of the Study:
- To compare the accuracy of four leading large language models (LLMs) on cornea and external disease multiple-choice questions (MCQs).
- To assess the safety of LLM responses by identifying potentially harmful incorrect answers.
- To establish benchmarks for LLM performance in ophthalmology.
Main Methods:
- Four LLMs (DeepSeek-V3.2, GPT-5.2, Gemini 3 Pro, Claude Opus 4.5) were tested on American Academy of Ophthalmology (AAO) Cornea/External MCQs.
- A standardized prompt was used for single-query analysis, with a secondary five-query retest.
- Accuracy was statistically compared, and potentially harmful wrong answers were independently coded.
Main Results:
- GPT-5.2, Gemini 3 Pro, and Claude Opus 4.5 achieved high accuracy (92-94%) on the AAO ONE dataset, not significantly different from each other.
- DeepSeek-V3.2 showed significantly lower accuracy (65.8%) in the AAO ONE dataset.
- In the AAO Basic and Clinical Science Course dataset, Gemini 3 Pro (94.7%) and Claude Opus 4.5 (84.2%) outperformed DeepSeek-V3.2 (57.9%), with GPT-5.2 at 76.3%.
Conclusions:
- Top-tier LLMs (GPT-5.2, Gemini 3 Pro, Claude Opus 4.5) exhibit comparable high accuracy on text-based cornea/external disease MCQs.
- DeepSeek-V3.2 significantly underperformed compared to the other evaluated models.
- While potentially harmful errors were infrequent, continued clinician oversight remains essential for safe LLM application in ophthalmology.
