Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Apr 14, 2026

Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization
05:49

Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization

Published on: February 23, 2024

1.7K

Digital guides in eye care: Comparing AI model accuracy and reliability.

Hakan Veli Savaş1, Osman Altay2

  • 1Department of Ophthalmology, Karakoçan State Hospital, Karakoçan, Elazığ, Turkey.

Digital Health
|April 13, 2026
PubMed
Summary

Four large language models (LLMs) were evaluated for ophthalmology patient education. Gemini and ChatGPT showed higher accuracy and reliability, while LLaMA performed poorly, with some models generating unsafe content.

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Intravitreal Nilotinib Reduces Fibroblast Growth Factor-2 Levels in a Rat Model of Dispase-Induced Intraocular Inflammation.

Journal of vitreoretinal diseases·2026
Same author

Assessing the therapeutic potential of astaxanthin in experimental proliferative vitreoretinopathy models.

European journal of ophthalmology·2025
Same author

A novel chaotic transient search optimization algorithm for global optimization, real-world engineering problems and feature selection.

PeerJ. Computer science·2023
See all related articles

Area of Science:

  • Ophthalmology
  • Artificial Intelligence
  • Medical Education

Background:

  • Large language models (LLMs) are increasingly used for patient education.
  • Evaluating the accuracy, reliability, and safety of LLMs in specialized medical fields like ophthalmology is crucial.
  • Patient education in ophthalmology requires precise and safe information delivery.

Purpose of the Study:

  • To comparatively evaluate the performance of four leading LLMs (ChatGPT, Gemini, Claude, LLaMA) for patient education in ophthalmology.
  • To assess LLM accuracy, reliability, and patient safety across various ophthalmic subspecialties.
  • To identify potential risks and benefits of using LLMs in ophthalmic patient communication.

Main Methods:

  • A cross-sectional evaluation involving 50 frequently asked patient questions across five ophthalmic subspecialties.
Keywords:
Ophthalmologyexpert evaluationlarge language modelspatient educationpatient safety

More Related Videos

Artificial Intelligence Approaches to Assessing Primary Cilia
08:58

Artificial Intelligence Approaches to Assessing Primary Cilia

Published on: May 1, 2021

4.3K

Related Experiment Videos

Last Updated: Apr 14, 2026

Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization
05:49

Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization

Published on: February 23, 2024

1.7K
Artificial Intelligence Approaches to Assessing Primary Cilia
08:58

Artificial Intelligence Approaches to Assessing Primary Cilia

Published on: May 1, 2021

4.3K
  • Text-only questions were submitted to ChatGPT o3 Mini High, Gemini 2.0 Pro, Claude-Sonnet 3.7, and LLaMA 3.1 405B.
  • Responses were independently assessed by five blinded ophthalmologists on accuracy, currency, clarity, and patient safety, with unsafe content categorized.
  • Main Results:

    • Significant performance variations were observed among the LLMs, with mean scores: Gemini (3.44), ChatGPT (2.99), Claude (2.48), and LLaMA (1.09).
    • Gemini generally outperformed other models, though ChatGPT and Claude showed strength in the retina subspecialty.
    • Potentially unsafe content was present in 9.5% of responses, with LLaMA exhibiting the highest proportion and Gemini the lowest.

    Conclusions:

    • LLMs offer potential for ophthalmology patient education, but performance varies by model and subspecialty.
    • Gemini 2.0 Pro and ChatGPT o3 Mini High demonstrated relatively higher accuracy and reliability in this evaluation.
    • Further clinical studies are necessary to determine the safe and effective integration of LLMs into ophthalmic practice, assessing patient comprehension and behavior.