Related Experiment Video
Updated: Aug 16, 2026

Binocular Dynamic Visual Acuity in Eyeglass-Corrected Myopic Patients
Published on: March 29, 2022
Evaluating multimodal large language models for differential diagnosis of high myopia versus high myopia with
Houfa Yin1,2, Tao Yu1,2, Wei Wu1,2
1Eye Center of Second Affiliated Hospital, School of Medicine, Zhejiang University, Hangzhou, China.
Background:
Distinguishing glaucomatous damage from structural and functional changes related to high myopia remains a persistent clinical challenge.
Objective:
To evaluate multimodal large language models (LLMs) for the differential diagnosis of high myopia alone versus high myopia with glaucoma under direct-answer and chain-of-thought (CoT) prompting, and to compare physician-rated response quality across models.
Methods:
We analyzed a fixed retrospective archive of 100 patients, including 50 patients with high myopia without glaucoma and 50 patients with high myopia and primary open-angle glaucoma (POAG). High myopia was defined according to the source clinical archive as spherical equivalent of -6.0 diopters or less and/or axial length of 26.0 mm or greater. Four multimodal LLMs were assessed in Direct and CoT modes. Two ophthalmologists, each with more than 10 years of clinical experience, independently rated each output on clinical logic, medical knowledge/conclusion accuracy, and evidence use/multimodal consistency; one rater was a chief physician and the other was an attending physician. Diagnostic correctness was compared using Cochran's Q and McNemar tests; physician ratings were compared using Friedman and Wilcoxon signed-rank tests; weighted Cohen's kappa and intraclass correlation coefficients (ICCs) were used to assess inter-rater agreement.
Results:
In Direct mode, case-level accuracies were 83.0% for Gemini 3 Flash Preview, 75.0% for Kimi-k2.5, 65.0% for GPT-5.4, and 49.0% for Qwen3.5-Plus. In CoT mode, case-level accuracies were 78.0% for Gemini 3 Flash Preview, 75.0% for Qwen3.5-Plus, 73.0% for GPT-5.4, and 69.0% for Kimi-k2.5. Cross-model differences were significant in Direct mode (Cochran's Q = 31.16, p < 0.001) but not in CoT mode (Q = 3.23, p = 0.357). Inter-rater agreement was strong, with weighted Cohen's kappa values of 0.80, 0.84, and 0.83 across the three physician-rated domains and ICC(3, k) = 0.92 for the combined total score.
Conclusion:
Multimodal LLMs showed variable performance on this retrospective benchmark for differentiating high myopia from high myopia with glaucoma, but CoT prompting did not consistently improve diagnostic accuracy. Physician ratings provided complementary evidence on clinical coherence and multimodal evidence use, supporting a broader benchmark-oriented evaluation framework for ophthalmic LLM applications.