Related Experiment Video
Updated: Jun 9, 2026

In vivo Structural Assessments of Ocular Disease in Rodent Models using Optical Coherence Tomography
Published on: July 24, 2020
Evaluating large language models vs residents in cataract and refractive surgery: comparative analysis using the
Avi Wallerstein1, Taanvee Ramnawaz, Mathieu Gauvin
1From the Department of Ophthalmology and Visual Sciences, McGill University, Montreal, Quebec, Canada (Wallerstein, Gauvin); LASIK MD, Montreal, Quebec, Canada (Wallerstein, Ramnawaz, Gauvin); School of Medicine, University of Montreal, Montreal, Quebec, Canada (Ramnawaz).
Purpose:
To assess the accuracy of leading large language models (LLMs) in answering cataract and refractive surgery questions, determine whether prompt complexity affects performance, and compare their accuracy with ophthalmology residents.
Setting:
Not applicable.
Design:
Randomized questionnaire-based study using 100 questions from the cataract and refractive surgery section of the Basic and Clinical Science Course (BCSC) Self-Assessment Program.
Methods:
5 LLMs (ChatGPT-4, ChatGPT-4o, Gemini, Gemini Advanced, and Copilot-Precise Mode) were tested. Data from 1983 BCSC Self-Assessment Program users served as a comparison. Each LLM underwent 2 sessions: one with a simple prompt and another with a contextualized prompt. Accuracy was defined as the proportion of correct answers.
Results:
Using the simple prompt, ChatGPT-4o achieved the highest accuracy at 84% (95% CI 77%-91%), followed by Gemini Advanced at 82%, Copilot at 78%, ChatGPT-4 at 77%, and Gemini at 62% ( P < .05). With the contextualized complex prompt, ChatGPT-4o again led (86%; 95% CI 79%-93%). Performance differences between simple and complex prompts were not statistically significant ( P > .05). Except for Gemini, all models' lower 95% CIs exceeded the 60% passing threshold. The mean resident score was 77% (95% CI 75%-79%). Only ChatGPT-4o significantly ( P = .04) outperformed residents, while Gemini Advanced and Copilot trended higher ( P ∼ .10).
Conclusions:
ChatGPT-4o consistently outperformed ChatGPT-4, Gemini, Gemini Advanced, and Copilot, and was the only model to significantly surpass residents. Prompt complexity did not affect LLM performance. All models except Gemini exceeded the 60% accuracy threshold, indicating the potential of LLMs as tools for knowledge assessment in refractive surgery.
More Related Videos
05:46Author Spotlight: Advancements in Refractive Surgical Correction for Presbyopia and Exploring Postoperative Visual Acuity
Published on: September 20, 2024
05:14Comparison of Agreement and Accuracy using Binocular Wavefront Optometer with Autorefractor and Phoropter
Published on: September 16, 2025
Related Concept Videos
Sleep Apnea
The condition is more prevalent among...
Cardiopulmonary Resuscitation II: ACLS Airway Management