Comparative Accuracy and Safety of 4 Large Language Models on Cornea and External Disease Multiple-Choice Questions

Bora Deniz Argon1, Şule Vildan Durmuş

  • 1Department of Ophthalmology, University of Health Sciences, Prof. Dr. Cemil Taşcıoğlu City Hospital, Şişli, Istanbul, Turkey.

Cornea
|June 16, 2026
PubMed
Summary

Four large language models were evaluated on ophthalmology questions. GPT-5.2, Gemini 3 Pro, and Claude Opus 4.5 demonstrated high accuracy, while DeepSeek-V3.2 performed significantly lower, indicating a need for continued clinician oversight.

Related Concept Videos