Related Experiment Video
Updated: Jun 19, 2026

In Vivo Vascular Injury Readouts in Mouse Retina to Promote Reproducibility
Published on: April 21, 2022
Accuracy and Reproducibility of Different Artificial Intelligence Chatbots' Responses to Patient-Based Vitreoretinal
Motasem Al-Latayfeh1,2, Abdelwahab Aleshawi3, Omar S El-Mulki4
1Department of Special Surgery, Faculty of Medicine, the Hashemite University, Zarqa, Jordan.
Generative artificial intelligence (AI) chatbots show promise for patient education in vitreoretinal care, with ChatGPT-5.o and DeepSeek R1 demonstrating high accuracy and reproducibility. However, performance varies by model and condition, necessitating careful clinical implementation.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Informatics
Background:
- Generative artificial intelligence (AI) chatbots are increasingly utilized by patients seeking health information.
- The reliability of AI chatbots in providing accurate information for complex ophthalmic conditions is not well-established.
- This study evaluates the performance of leading AI chatbots in responding to patient-centered vitreoretinal queries.
Purpose of the Study:
- To compare the accuracy, comprehensiveness, and reproducibility of five AI chatbots in answering patient questions about vitreoretinal conditions.
- To assess the suitability of AI chatbots as patient education tools in ophthalmology.
Main Methods:
- 135 patient-centered vitreoretinal questions from the American Academy of Ophthalmology database were used.
- Five AI chatbots (ChatGPT-5.o, DeepSeek R1, Meta AI, Grok 3.0, Google Gemini 2.5 Pro) were queried twice under standardized conditions.
- Two independent vitreoretinal ophthalmologists evaluated response accuracy and reproducibility.
Main Results:
- ChatGPT-5.o achieved the highest accuracy (94%) and high reproducibility (96.3%).
- DeepSeek R1 showed the greatest reproducibility (98.5%) with high accuracy (92.6%).
- Grok 3.0 and Google Gemini 2.5 Pro exhibited lower accuracy and reproducibility, while Meta AI performed moderately.
- Vitrectomy and age-related macular degeneration questions received higher accuracy scores compared to diabetic retinopathy.
Conclusions:
- ChatGPT-5.o and DeepSeek R1 demonstrate accuracy and reproducibility comparable to clinical standards, suggesting potential as patient education resources.
- Variability in performance across different AI models and ophthalmic conditions underscores the need for cautious adoption.
- Continued optimization is crucial to ensure AI chatbots deliver safe and reliable information for vitreoretinal care.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025