Related Experiment Videos
Large Language Models Approximate Inter-Expert Agreement in Glaucoma Suspect and Glaucoma Classification from
Ryan S Shean1, Jayanth Kumar Mallapu1, Tathya Shah1
1Keck School of Medicine, University of Southern California, Los Angeles, California.
Objective:
To assess whether publicly available large language models (LLMs), when provided cross-sectional multimodal clinical inputs, can classify glaucoma versus glaucoma suspect with diagnostic agreement comparable to fellowship-trained glaucoma specialists when benchmarked against consensus-derived reference standards.
Design:
Observational cross-sectional study.
Subjects And Controls:
A total of 230 eyes from 131 consecutive participants evaluated at a tertiary academic glaucoma referral center in the United States between 2016 and 2022 were included. The eye was the unit of analysis. No separate external control group was used; comparisons were made against peer-derived consensus reference standards.
Methods:
Seven LLM configurations (GPT-5 Pro, GPT-5.2, Gemini 3 Pro under two prompts, and Grok 2.2 under 1 prompt) were tested without task-specific training using multimodal inputs including age, sex, race, visual acuity, intraocular pressure, fundus photographs, OCT retinal nerve fiber layer reports, and visual field reports. Four peer-derived reference standards were constructed using a leave-one-out majority consensus approach among four fellowship-trained glaucoma specialists who independently graded the complete data.
Main Outcomes And Measures:
Diagnostic agreement for classification as glaucoma suspect or glaucoma was assessed using accuracy, sensitivity, specificity, F1 score, and Cohen's κ.
Results:
Among 230 eyes of 131 patients (72 women [55%]; 49 Asian, 10 Black, 56 Caucasian, 50 Hispanic, and 65 other), mean (standard deviation) age was 67.4 (13.8) years, and 43.5% to 56.5% of eyes were classified as glaucoma. Glaucoma specialist accuracy ranged from 71.3% to 83.9% (κ = 0.46-0.68). GPT-5 Pro (long prompt) achieved accuracies of 80.4% to 85.7% (κ = 0.61-0.71), and Gemini 3 Pro (long prompt) achieved accuracies of 80.9% to 84.3% (κ = 0.62-0.68), each achieving the highest accuracy in two reference sets. GPT-5.2 demonstrated intermediate performance; Grok 2.2 performed near chance. Agreement was highest for moderate-to-severe glaucoma and lower for glaucoma suspect and mild glaucoma.
Conclusions:
Publicly available multimodal LLMs achieved diagnostic agreement comparable to inter-expert agreement among glaucoma specialists without task-specific training. These findings support further investigation of LLMs as scalable, standardized clinical decision-support tools in glaucoma care.
Financial Disclosures:
Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.