Related Experiment Video
Updated: Mar 31, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessment of Correctness, Content Omission, and Risk of Harm in Large Language Model Responses to Ophthalmology
Jacqueline L Chen1, Amanda J Lu2, Rohan Verma3
1Sidney Kimmel Medical College, Thomas Jefferson University, Philadelphia, Pennsylvania.
Purpose:
To evaluate the accuracy and prose responses of 2 large language models (LLMs) to ophthalmology continuing medical education questions.
Design:
Question prompts and multiple choice (MC) answer options were input into the 2 LLMs, and responses were analyzed for accuracy and assessed for evidence of correctness, completeness, bias, and potential harm using a previously reported standardized rubric.
Subjects:
Basic and Clinical Science Course questions and MC answer options from the American Academy of Ophthalmology question bank were used as inputs into the 2 LLMs (ChatGPT-4 and Google Vertex's Gemini Pro 1.5).
Methods:
The MC responses were assessed for accuracy in comparison to the question bank's designated corrected answer. The free-text prose responses from the 2 LLMs were assessed by 3 board-certified ophthalmologists.
Main Outcome Measures:
Accuracy and assessment of correct and incorrect reasoning, inappropriate content, missing content, possibility of bias, or possibility of harm.
Results:
The MC accuracy rates of ChatGPT-4 and Gemini Pro 1.5 were 82.5% (99/120) and 49.2% (59/120) (P < 0.05), respectively. Though there was high evidence of correct reasoning in the prose responses (92% and 88% for ChatGPT-4 and Gemini Pro 1.5, respectively), there was also evidence of incorrect reasoning (42% and 58%), inappropriate content (29% and 36%), missing content (42% and 30%), and possibility of physical or emotional harm (36% and 44%).
Conclusions:
Though ChatGPT-4 was able to perform well in MC accuracy, both LLMs contained inaccuracies, missing content, and material that could lead to harm in their prose responses. Our findings suggest that provider-guided auditing in ophthalmology is required before the use of the technology in direct patient-facing settings.
Financial Disclosures:
Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
Related Concept Videos
Glaucoma: Overview
Accessory Structures of the Eye
Open Angle Glaucoma: Treatment
Drugs such as carbonic anhydrase inhibitors, α2- and...
Angle Closure Glaucoma: Treatment
Assessment of Airway, Skin Color, and Use of Accessory Muscles
Introduction
The initial evaluation of a patient's respiratory system...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...

