Related Experiment Video
Updated: Jun 16, 2026

Assessing Early Stage Open-Angle Glaucoma in Patients by Isolated-Check Visual Evoked Potential
Published on: May 25, 2020
Assessing GPT-4o in cataract surgery decision-making: appropriateness, consistency, and clinical implications
Yan Liu1,2,3,4, Chaozhong Zhang5, Yuanping Yuan6
1Eye Institute and Department of Ophthalmology, Eye and ENT Hospital, Fudan University, Shanghai, China.
Purpose:
To evaluate the application of large language models (LLMs) in clinical decision-making for cataract surgery using real-world patient data.
Methods:
This retrospective clinical observational study included 74 cataract patients (146 eyes) who underwent phacoemulsification and intraocular lens (IOL) implantation. Preoperative datasets, including clinical findings from ophthalmic imaging, biometric parameters, systemic histories, and lifestyle questionnaires, were processed via GPT-4o. The model was prompted to simulate an experienced cataract surgeon. GPT-4o's IOL recommendations were evaluated for appropriateness with an adjudicated expert consensus and compared against selections made by high- and low-volume cataract surgeons. Appropriateness rates, agreement (Kappa value), and consistency (intraclass correlation coefficient, ICC) were analyzed.
Results:
The overall appropriateness rate of GPT-4o for IOL selection was significantly lower than that of human surgeons (37.84% vs. 72.14%, P = 0.001). Subgroup analysis revealed that appropriateness varied by IOL category (P < 0.001), with the highest alignment observed for monofocal IOLs (90.0%) and the lowest for trifocal toric IOLs (9.09%). The presence of ocular or systemic comorbidities was associated with a slight reduction in appropriateness. While agreement analysis indicated poor inter-rater agreement (Kappa < 0.4), GPT-4o demonstrated a slightly higher ICC (0.378) than human surgeons (0.359). However, repeated testing showed significant instability in the model's responses (P = 0.026).
Conclusion:
GPT-4o demonstrated moderate appropriateness for routine monofocal IOL selection but showed limited appropriateness and consistency in complex, customized scenarios. These findings suggest GPT-4o could assist routine decision-making in cataract surgery, while highlighting the need for further domain-specific fine-tuning and multimodal optimization for advanced cases.
Related Concept Videos
Glaucoma: Overview
Angle Closure Glaucoma: Treatment
Open Angle Glaucoma: Treatment
Drugs such as carbonic anhydrase inhibitors, α2- and...