Related Experiment Video
Updated: Jun 27, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating the Use of Large Language Models in Improving the Readability of Online Patient Education Materials for
Nikhil Sriram1, Rishi Jain1, Mehul Mittal1
1Department of Neurological Surgery, Northwestern University Feinberg School of Medicine, Chicago, IL 60611, USA.
Abstract:
Objective: Online patient education materials (OPEMs) are important resources for patients seeking health information. While the National Institutes of Health (NIH) and American Medical Association (AMA) recommend a sixth-grade readability level for OPEMs, commonly available material often exceeds such criteria. Large language models (LLMs), such as ChatGPT and Gemini, have emerged as tools for health education with potential applications in simplification of health material. This study assesses the utility of ChatGPT and Gemini in enhancing the readability of OPEMs for peripheral nerve surgeries. Methods: Eleven common peripheral nerve surgeries were used as online search terms. The first 20 unique search results were assessed; results were excluded if they did not include patient-facing material. ChatGPT and Gemini were instructed to rewrite the text of the OPEM at or below a sixth-grade reading level. Readability metrics were calculated for original OPEMs, alongside ChatGPT and Gemini rewrites. LLM responses were reviewed for accuracy/quality (five-point scale) and comprehensiveness (three-point scale) using predefined criteria. Results: A total of 220 websites were assessed. In total, 155 OPEMs met the inclusion criteria; 65 websites were excluded because they were academic journal articles or other provider-facing materials. The average Flesch-Kincaid grade level (FKGL) of OPEMs was 11.3, significantly greater than the NIH/AMA-sixth grade recommendations (p < 0.001). The average FKGL of ChatGPT rewrites was significantly lower than that of OPEMs (11.3 vs. 7.5, p < 0.001), as was the average FKGL of Gemini rewrites (11.3 vs. 5.6, p < 0.001). ChatGPT rewrites were of higher accuracy/quality (4.5/5.0 vs. 4.0/5.0, p < 0.001) and comprehensiveness (2.0/3.0 vs. 1.0/3.0, p < 0.001) relative to Gemini rewrites. Conclusions: The readability of online patient education materials for peripheral nerve surgery significantly exceeded NIH/AMA recommendations. ChatGPT and Gemini were able to significantly simplify the reading level of these OPEMs. LLMs may serve as tools to improve the readability of peripheral nerve surgery OPEMs.
Related Concept Videos
Introduction to Language of Pathophysiology l
Introduction to Language of Pathophysiology ll
