Related Experiment Video
Updated: Aug 6, 2026

05:14
Comparison of Agreement and Accuracy using Binocular Wavefront Optometer with Autorefractor and Phoropter
Published on: September 16, 2025
Evaluation of the large language models Chatbot's responses to frequently asked queries on refractive errors
Sare Safi1,2, Amir Golmakani2, Saeed Rahmani2
1Ophthalmic Epidemiology Research Center, Research Institute for Ophthalmology and Vision Science, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
Health Informatics Journal
|July 20, 2026
Summary
Four Large Language Model (LLM) chatbots demonstrated comparable accuracy in answering refractive error questions. ChatGPT offered the most accessible responses, with all chatbots exceeding public health readability standards.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Natural Language Processing
Background:
- Refractive errors are common vision impairments.
- Large Language Models (LLMs) are increasingly used for health information.
- Assessing the reliability of LLM-generated health information is crucial.
Purpose of the Study:
- To evaluate the accuracy, comprehensiveness, and reliability of four LLM chatbots (Copilot, Perplexity, Gemini, ChatGPT) in responding to frequently asked questions about refractive errors.
- To compare the readability and response similarity across different LLM chatbots.
Main Methods:
- Forty-four refractive error questions were input into four LLM chatbots.
- Expert evaluation using a three-point accuracy scale assessed response quality.
- Readability indices and Sentence-Bidirectional Encoder Representations from Transformers (SBERT) measured text characteristics.
- Inter-rater agreement was quantified using Gwet's Agreement Coefficient 1 (AC1).
Main Results:
- Near-perfect inter-rater agreement (Gwet AC1: 0.87) was achieved.
- All chatbots scored above 85% in accuracy, with no significant differences (P=0.168).
- Perplexity showed the highest mean comprehensiveness (P=0.028), while ChatGPT provided more accessible output.
- Significant differences in readability (P<0.001) were noted among chatbots.
Conclusions:
- LLM chatbots exhibit comparable accuracy for refractive error information.
- All evaluated chatbots produce responses exceeding recommended public health readability thresholds.
- ChatGPT's output was found to be the most accessible among the tested LLMs.

