Related Experiment Video
Updated: Aug 5, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Using Large Language Models to Generate Retina Patient Education Material: A Comparative Analysis With American
Khaldon F Abbas1, Monali S Malvankar-Mehta1,2,3, Verena Juncal1,2
1Ivey Eye Institute, St. Joseph's Health Care Centre, London, ON, Canada.
Purpose:
To evaluate the readability, quality, and misinformation of patient education materials generated by large language models, including ChatGPT-4o (OpenAI), Gemini 1.5 Pro (Google), and Copilot Pro (Microsoft), compared with American Society of Retina Specialists (ASRS) brochures for retinal diseases.
Methods:
A cross-sectional comparative analysis was performed by generating patient education materials on 3 retinal conditions: retinal detachment, diabetic retinopathy, and age-related macular degeneration. Materials were created using a general prompt (prompt A) and a prompt specifying a sixth-grade readability level (prompt B). Readability was evaluated using 6 validated metrics. Quality was assessed through DISCERN and the Patient Education Materials Assessment Tool. Misinformation was graded using a 5-point Likert scale. Assessments were performed independently by 2 masked retina specialists.
Results:
Average readability of Gemini (11.65; P = .005) and Copilot (11.23; P = .003) materials was significantly better than that of ASRS materials (14.17), whereas ChatGPT showed no significant difference (12.85; P = .06). ChatGPT's average readability was significantly lower compared with Gemini (12.85 vs 11.65; P = .01) and Copilot (12.85 vs 11.23; P < .001). Prompt B significantly improved readability across all large language models relative to ASRS but still exceeded the sixth-grade readability level. DISCERN scores were comparable across groups. ASRS materials had an understandability score of 74.37%, which was significantly lower than ChatGPT (94.44%; P = .02) and Gemini 1.5 (95.83%, P = .02) scores. No significant differences were observed for actionability or misinformation. Readability showed no significant correlation with quality or misinformation (P > .05).
Conclusions:
Large language models, when appropriately prompted, can generate retina-related patient education material with superior readability compared with existing ASRS brochures, while maintaining comparable quality and accuracy. Large language models represent a promising approach for addressing literacy barriers, though expert oversight remains essential.