Related Experiment Video
Updated: Aug 6, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessing a large language model for glaucoma knowledge: ChatGPT-5 versus residents
Mauro Gobira1,2, Rodrigo Moreira1, Flavio J L Galhardo Carvalho Filho1
1Ophthalmology Department, Instituto da Visão, São Paulo, SP, Brazil.
Arquivos Brasileiros De Oftalmologia
|July 22, 2026
Summary
ChatGPT-5, a large language model, significantly outperformed ophthalmology residents on glaucoma knowledge tests. While advanced in many areas, residents showed strength in questions the AI missed, highlighting AI's potential and limitations in medical education.
Area of Science:
- Ophthalmology
- Artificial Intelligence in Medicine
- Medical Education Technology
Background:
- Large language models (LLMs) are increasingly integrated into various professional fields.
- Assessing the performance of LLMs against human experts is crucial for understanding their capabilities and limitations.
- Glaucoma education and assessment tools are vital for training ophthalmologists.
Purpose of the Study:
- To compare the performance of a contemporary LLM, ChatGPT-5, against ophthalmology residents.
- To evaluate the accuracy of ChatGPT-5 on standardized glaucoma multiple-choice questions.
- To identify areas where LLMs excel or fall short compared to resident knowledge.
Main Methods:
- A cross-sectional study utilized 189 text-only glaucoma multiple-choice questions from the Cybersight question bank.
- ChatGPT-5 was tested under standardized, text-only conditions.
- Six ophthalmology residents (PGY1-3) answered the same questions under supervision, with accuracy compared using McNemar's exact test.
Main Results:
- ChatGPT-5 achieved 86.8% accuracy, significantly outperforming residents (62.9% overall accuracy).
- ChatGPT-5 outperformed all residents in head-to-head comparisons (ORs 1.84-13.15, p≤0.023).
- Residents were more successful on 9.0% of items that ChatGPT-5 answered incorrectly.
Conclusions:
- ChatGPT-5 demonstrates potential as a valuable tool for ophthalmology education and assessment.
- Current LLM performance is limited by text-only data and specific question banks.
- Further research with multimodal data and larger cohorts is needed to confirm generalizability and clinical utility.
