Related Experiment Videos
Comparative evaluation of large language models and clinicians in real-world glaucoma clinical reasoning
Houfa Yin1,2, Lixia Shen1,2, Haiyan Cai1,2
1Eye Center of Second Affiliated Hospital, School of Medicine, Zhejiang University, Hangzhou, China.
Purpose:
Clinical decision-making in glaucoma is complex and requires integration of heterogeneous information, including patient history, examination findings, and risk stratification. While artificial intelligence (AI) has shown strong performance in image-based ophthalmic tasks, its capability in specialty-specific clinical reasoning remains insufficiently explored.
Methods:
Performance was evaluated by glaucoma specialists using a predefined rubric across three clinically oriented domains: medical accuracy (40%), key-point recall (30%), and logical completeness (30%). The weighted composite score was used as a descriptive summary of case-based reasoning quality.
Results:
AI models showed structured clinical reasoning performance in this case-based dataset, with weighted mean scores overlapping with those of attending ophthalmologists and exceeding those of some lower-performing trainees. These findings should be interpreted as exploratory performance patterns rather than evidence of equivalence. Inter-individual variability was substantial among human clinicians, particularly residents. AI systems often included safety-critical diagnostic and management elements, while the best-performing human clinician achieved the highest individual score overall.
Conclusion:
In this limited 34-case evaluation, large language model-based AI systems produced structured glaucoma-related reasoning with performance that overlapped with attending ophthalmologists but did not establish clinical equivalence. These systems require specialist oversight and further validation before clinical use, but may have potential as supervised decision-support and educational tools.