Related Experiment Video
Updated: Jun 9, 2026

In vivo Structural Assessments of Ocular Disease in Rodent Models using Optical Coherence Tomography
Published on: July 24, 2020
Evaluating large language models vs residents in cataract and refractive surgery: comparative analysis using the
Avi Wallerstein1, Taanvee Ramnawaz, Mathieu Gauvin
1From the Department of Ophthalmology and Visual Sciences, McGill University, Montreal, Quebec, Canada (Wallerstein, Gauvin); LASIK MD, Montreal, Quebec, Canada (Wallerstein, Ramnawaz, Gauvin); School of Medicine, University of Montreal, Montreal, Quebec, Canada (Ramnawaz).
ChatGPT-4o demonstrated superior accuracy in answering cataract and refractive surgery questions, significantly outperforming ophthalmology residents and other large language models (LLMs). Prompt complexity did not impact LLM performance.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Education
Background:
- Large language models (LLMs) are increasingly used in various fields, including medicine.
- Assessing the accuracy of LLMs in specialized medical domains like ophthalmology is crucial.
- Cataract and refractive surgery knowledge is essential for ophthalmologists.
Purpose of the Study:
- To evaluate the accuracy of leading LLMs in answering questions related to cataract and refractive surgery.
- To determine if prompt complexity influences LLM performance in this domain.
- To compare LLM accuracy against ophthalmology residents' performance.
Main Methods:
- A randomized questionnaire-based study was conducted using 100 questions from the cataract & refractive surgery section of the BCSC Self-Assessment Program.
- Five LLMs (ChatGPT-4, ChatGPT-4o, Gemini, Gemini Advanced, and Copilot-Precise Mode) were tested.
- LLMs were evaluated using both simple and contextualized prompts, with accuracy defined as the proportion of correct answers. Data from 1,983 BCSC users served as a comparison.
Main Results:
- ChatGPT-4o achieved the highest accuracy (84% with simple prompts, 86% with complex prompts).
- Performance differences between simple and complex prompts were not statistically significant for any LLM.
- ChatGPT-4o was the only model to significantly outperform ophthalmology residents (mean score 77%), while Gemini Advanced and Copilot showed trending higher performance.
Conclusions:
- ChatGPT-4o demonstrated superior and consistent performance across prompt types and significantly surpassed ophthalmology residents.
- Prompt complexity did not significantly alter LLM accuracy in this study.
- LLMs, particularly ChatGPT-4o, show potential as valuable tools for knowledge assessment in refractive surgery, with most models exceeding a 60% accuracy threshold.
More Related Videos
05:46Author Spotlight: Advancements in Refractive Surgical Correction for Presbyopia and Exploring Postoperative Visual Acuity
Published on: September 20, 2024
05:14Comparison of Agreement and Accuracy using Binocular Wavefront Optometer with Autorefractor and Phoropter
Published on: September 16, 2025
Related Concept Videos
Sleep Apnea
The condition is more prevalent among...
Cardiopulmonary Resuscitation II: ACLS Airway Management