Related Experiment Video
Updated: Mar 27, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.3K
Evaluating Three Large Language Models in Ophthalmology Education: A Comparative Study of Accuracy and Readability
Semih Çakmak1, Arzu Karakiraz2, Işılsu Ezgi Uluisik1
1Department of Ophthalmology, İstanbul Faculty of Medicine, İstanbul University, İstanbul, Türkiye.
Ophthalmic Epidemiology
|March 25, 2026
Summary
Microsoft Copilot demonstrated superior accuracy and readability compared to ChatGPT and Gemini for ophthalmology exam questions. AI chatbot performance varies, impacting their educational utility.
Area of Science:
- Medical Education
- Artificial Intelligence
- Ophthalmology
Background:
- Large language models (LLMs) are increasingly used in education.
- Evaluating the performance of different LLMs is crucial for their effective implementation.
Purpose of the Study:
- To compare the accuracy and readability of responses generated by ChatGPT-4.0 Mini, Gemini 1.5 Flash, and Microsoft Copilot.
- To assess the suitability of AI-generated content for medical school ophthalmology exams.
Main Methods:
- 442 multiple-choice ophthalmology questions were submitted to each chatbot.
- Response accuracy was determined using an answer key.
- Readability was assessed using Flesch-Kincaid Grade Level (FKGL), Flesch Reading Ease (FRE), and Simple Measure of Gobbledygook (SMOG) indices.
Main Results:
- Microsoft Copilot achieved the highest accuracy rate (89.4%), followed by ChatGPT (84.2%) and Gemini (76.7%).
- Significant differences in readability were observed (p < 0.001), with Copilot showing the lowest linguistic complexity and ChatGPT the highest.
- All chatbot responses generally required a high school to early college reading level.
Conclusions:
- LLM-based chatbots exhibit variable performance in answering medical-level ophthalmology questions.
- Microsoft Copilot outperformed ChatGPT and Gemini in both accuracy and readability.
- The choice of AI model can significantly impact the usefulness of AI-generated content in educational contexts.

