Related Experiment Video
Updated: Aug 21, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Ability of Large Language Models to Answer Patients' Questions and Generate Educational Materials for Uncommon
Samuel A Cohen1, Prashant D Tailor1, Adrian Au1
1Department of Ophthalmology, UCLA Jules Stein Eye Institute, University of California Los Angeles, Los Angeles, CA, USA.
Purpose:
To assess the ability of large language models to accurately and comprehensively respond to frequently asked questions and generate patient education materials related to uncommon retinal conditions.
Methods:
A total of 50 frequently asked questions related to 10 uncommon retinal conditions were input into 3 large language models: ChatGPT-4o1, Google Gemini 2.0 Flash, and Microsoft Copilot (updated January 7, 2025). The accuracy and completeness of responses to frequently asked questions were evaluated by retina specialists using a Likert scale ranging from 1 (very inaccurate/not at all complete) to 5 (very accurate/completely complete), while readability was assessed using validated indices. The large language models were also instructed to generate patient education materials at specific reading levels, which were again evaluated for accuracy, completeness, and readability.
Results:
Responses to patient frequently asked questions were written at mean grade levels of 15.3 ± 1.4 (ChatGPT), 15.7 ± 1.9 (Gemini), and 14.6 ± 1.3 (Copilot), respectively (P = .02). Mean accuracy scores were 4.02 ± 0.7, 4.16 ± 0.7, and 3.81 ± 0.8. Accuracy scores for large language models-generated patient education materials were 4.70 ± 0.7 (ChatGPT), 4.85 ± 0.4 (Gemini), and 4.30 ± 0.8 (Copilot). When instructed to revise educational materials to improve understandability, all large language models significantly reduced reading levels by 4.8 (ChatGPT), 6.9 (Gemini), and 3.8 (Copilot) grade levels (P < .001) without compromising accuracy.
Conclusions:
Large language models can accurately respond to patients' frequently asked questions related to uncommon retinal conditions. Furthermore, large language models can effectively improve the readability of existing education materials for patients with varying levels of health literacy. A deeper understanding of large language model applications may facilitate their integration into clinical practice.