Related Experiment Video
Updated: Jan 7, 2026

05:14
Comparison of Agreement and Accuracy using Binocular Wavefront Optometer with Autorefractor and Phoropter
Published on: September 16, 2025
533
Comparative Performance Evaluation of Large Language Models and Human Teachers in Answering Optometry Questions from
Zijing Huang1, Tian Lin1,2, Huini Lin1
1Joint Shantou International Eye Center of Shantou University and The Chinese University of Hong Kong, Shantou, China.
Journal of Medical Education and Curricular Development
|January 5, 2026
Summary
Large language models (LLMs) show strong potential in optometry education, outperforming human teachers in answering student questions for accuracy and completeness. While LLMs provide detailed information, students still find human answers more comprehensible and are hesitant to replace educators.
Area of Science:
- Optometry Education
- Artificial Intelligence in Medicine
- Medical Training
Background:
- Medical undergraduate students often have complex questions requiring expert knowledge.
- Evaluating the efficacy of emerging technologies like large language models (LLMs) in specialized educational fields is crucial.
- Traditional teaching methods face challenges in providing immediate and comprehensive answers to diverse student queries.
Purpose of the Study:
- To compare the performance of five large language models (LLMs) against human teachers in answering optometry-related questions.
- To assess the accuracy, completeness, comprehensibility, and overall quality of answers provided by LLMs and human educators.
- To gauge student satisfaction and perspectives on integrating LLMs into optometry education.
Main Methods:
- A prospective, comparative study involving 108 questions from 30 medical undergraduate students.
- Answers were generated by human teachers and five LLMs (Mistral-7B, Llama-2-13B, Claude-3, Gemini-1.0 pro, GPT-4.0).
- Optometry experts blindly evaluated answers on a 5-point scale; students completed satisfaction questionnaires.
Main Results:
- LLMs provided faster and more extensive answers than humans (P < .001).
- Online LLMs (GPT-4.0, Claude-3, Gemini-1.0 pro) significantly outperformed human teachers and local LLMs in overall performance (P < .001).
- GPT-4.0 led in accuracy and completeness; Claude-3 excelled in comprehensibility and quality. Students found LLM information more comprehensive but human answers easier to understand.
Conclusions:
- Large language models (LLMs) demonstrate significant potential as supplementary tools in optometry education.
- LLMs can effectively address student queries, offering detailed and rapid responses.
- While beneficial, LLMs are not yet favored to replace human teachers entirely, highlighting the need for blended learning approaches.

