Related Experiment Video
Updated: Apr 14, 2026

05:49
Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization
Published on: February 23, 2024
1.7K
Digital guides in eye care: Comparing AI model accuracy and reliability.
Hakan Veli Savaş1, Osman Altay2
1Department of Ophthalmology, Karakoçan State Hospital, Karakoçan, Elazığ, Turkey.
Digital Health
|April 13, 2026
Summary
Four large language models (LLMs) were evaluated for ophthalmology patient education. Gemini and ChatGPT showed higher accuracy and reliability, while LLaMA performed poorly, with some models generating unsafe content.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Education
Background:
- Large language models (LLMs) are increasingly used for patient education.
- Evaluating the accuracy, reliability, and safety of LLMs in specialized medical fields like ophthalmology is crucial.
- Patient education in ophthalmology requires precise and safe information delivery.
Purpose of the Study:
- To comparatively evaluate the performance of four leading LLMs (ChatGPT, Gemini, Claude, LLaMA) for patient education in ophthalmology.
- To assess LLM accuracy, reliability, and patient safety across various ophthalmic subspecialties.
- To identify potential risks and benefits of using LLMs in ophthalmic patient communication.
Main Methods:
- A cross-sectional evaluation involving 50 frequently asked patient questions across five ophthalmic subspecialties.
- Text-only questions were submitted to ChatGPT o3 Mini High, Gemini 2.0 Pro, Claude-Sonnet 3.7, and LLaMA 3.1 405B.
- Responses were independently assessed by five blinded ophthalmologists on accuracy, currency, clarity, and patient safety, with unsafe content categorized.
Main Results:
- Significant performance variations were observed among the LLMs, with mean scores: Gemini (3.44), ChatGPT (2.99), Claude (2.48), and LLaMA (1.09).
- Gemini generally outperformed other models, though ChatGPT and Claude showed strength in the retina subspecialty.
- Potentially unsafe content was present in 9.5% of responses, with LLaMA exhibiting the highest proportion and Gemini the lowest.
Conclusions:
- LLMs offer potential for ophthalmology patient education, but performance varies by model and subspecialty.
- Gemini 2.0 Pro and ChatGPT o3 Mini High demonstrated relatively higher accuracy and reliability in this evaluation.
- Further clinical studies are necessary to determine the safe and effective integration of LLMs into ophthalmic practice, assessing patient comprehension and behavior.
