Related Experiment Video
Updated: Sep 18, 2025

Author Spotlight: Simple and Efficient Neural Retina Organoid Production for Disease Modeling
Published on: December 22, 2023
Implementing Generative AI to Enhance Patient Education on Retinopathy of Prematurity
Qais A Dihan1,2, Andrew D Brown3, Ana T Zaldivar4
1Chicago Medical School, Rosalind Franklin University of Medicine and Science, North Chicago, Illinois.
Purpose:
To evaluate the efficacy of large language models (LLMs) in generating patient education materials (PEMs) on retinopathy of prematurity (ROP).
Methods:
ChatGPT-3.5 (OpenAI), ChatGPT-4 (OpenAI), and Gemini (Google AI) were compared on three separate prompts. Prompt A requested that each LLM generate a novel PEM on ROP. Prompt B requested generated PEMs at the 6th-grade reading level using the validated Simple Measure of Gobbledygook (SMOG) readability formula. Prompt C requested LLMs improve the readability of existing, human-written PEMs to a 6th-grade reading level. PEMs inserted into Prompt C were sourced through a Google search of "retinopathy of prematurity." Each PEM was analyzed for readability (SMOG, Flesch-Kincaid Grade Level [FKGL]), quality (Patient Education Materials Assessment Tool [PEMAT], DISCERN), and accuracy (Likert Misinformation Scale).
Results:
LLM-generated PEMs were of high quality (median DISCERN = 4), understandable (PEMAT-U ≥ 70%), and accurate (Likert = 1). Prompt B generated more readable PEMs than Prompt A (P < .001). ChatGPT-4 and Gemini rewrote PEMs (Prompt C) from a baseline readability level (FKGL: 8.8 ± 1.9, SMOG: 8.6 ± 1.5) to the targeted 6th-grade reading level. Only ChatGPT-4 rewrites maintained high quality and reliability (median DISCERN = 4).
Conclusions:
LLMs, particularly ChatGPT-4, can serve as strong supplementary tools to automate the process of generating readable and high-quality PEMs for parents on ROP.

