Related Experiment Video
Updated: Apr 6, 2026

Binocular Dynamic Visual Acuity in Eyeglass-Corrected Myopic Patients
Published on: March 29, 2022
When Large Language Models (LLMs) walk into a Bachelor's in optometry examination: Comparing the performances of LLMs
Atul Arora1,2, Uday Pratap Singh Parmar3, Anjali1
1Advanced Eye Centre, Post Graduate Institute of Medical Education and Research, Chandigarh, India.
Purpose:
To evaluate the performance of Large Language Models (LLMs) on optometry examination questions and compare their accuracy and readability with Bachelor of Optometry students.
Methods:
A cross-sectional comparative study was conducted using the publicly available, free versions of five LLM models from four platforms (ChatGPT 3.5, ChatGPT 4o, Gemini, CoPilot, and DeepSeek) and a group of 15 third- and fourth-year optometry students. Two sets of multiple-choice questions (20 theoretical and 20 clinical) were administered to both the students and the LLMs. Theoretical questions covered core optometric knowledge, while clinical questions simulated real-life patient scenarios. Responses were graded by senior ophthalmologists for accuracy, and readability was assessed via readable.com using four indices, including Flesch-Kincaid Grade Level, Flesch Reading Ease Score, Coleman Liau Score, and Simple Measure of Gobbledygook (SMOG) Index.
Results:
The overall scores of the optometry students (28.13 ± 3.33) were comparable to those of the LLMs (29 ± 4.41). In theoretical questions, LLMs (15.40 ± 1.82) performed at par with the students (14.07 ± 2.21), with DeepSeek and CoPilot outperforming students (scoring 17 each). However, in clinical questions, the students performed better, highlighting the limitations of LLMs in context-specific reasoning. Pairwise comparisons of the readability analysis revealed that Gemini and DeepSeek provided significantly most readable explanations, while ChatGPT 3.5 produced the most complex responses. Across models, readability varied for Flesch-Kincaid grade level (P = 0.0213), Flesch Reading Ease Score (P = 0.0014), and SMOG (P = 0.0412), with a nonsignificant trend for Coleman Liau Score (P = 0.0529).
Conclusion:
LLMs show reasonable accuracy, matching students in theoretical performance but underperforming in clinical reasoning. Gemini and DeepSeek offer superior readability, highlighting their promise as educational tools. Future research should focus on integrating LLMs into curricula while balancing them with hands-on clinical education.
More Related Videos
06:19Comparison of Three Clinical Stereoscopic Methods for Measuring Binocular Visual Function During Amblyopic Treatment in Unilateral Amblyopia
Published on: September 27, 2024
05:14Comparison of Agreement and Accuracy using Binocular Wavefront Optometer with Autorefractor and Phoropter
Published on: September 16, 2025