Related Experiment Video
Updated: Jun 6, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparing the Readability and Usability of Patient Education Materials Generated by Different Large Language Models:
Alfonsus Adrian H Harsono1, Wendelyn M Oslock2, Bethany Brock3
1Division of Gastrointestinal Surgery, Department of Surgery, University of Alabama at Birmingham, Birmingham, Alabama.
Large language models (LLMs) can generate patient education materials (PEMs) to improve surgical care understanding. While Gemini improved readability, existing materials remain superior in usability and actionability.
Area of Science:
- Medical Informatics
- Health Literacy Research
- Artificial Intelligence in Healthcare
Background:
- Low health literacy presents significant barriers for patients navigating surgical care, contributing to health disparities.
- Large language models (LLMs) offer a potential scalable solution for enhancing patient education materials (PEMs) and improving patient comprehension.
- This study evaluates the readability and usability of PEMs generated by publicly available LLMs.
Purpose of the Study:
- To assess the readability of de novo patient education materials (PEMs) generated by three distinct large language models (LLMs).
- To evaluate the understandability and actionability of LLM-generated PEMs using the Patient Education Material Assessment Tool (PEMAT).
- To compare the performance of LLM-generated PEMs against existing baseline materials from an academic health center.
Main Methods:
- Existing colorectal surgical patient education materials (PEMs) were identified and used as a basis for generating new materials via LLMs (ChatGPT, Copilot, Gemini).
- Readability was assessed using Flesch-Kincaid reading ease, Flesch-Kincaid grade level, and modified grade-level scores.
- Usability was evaluated through understandability and actionability metrics using the Patient Education Material Assessment Tool (PEMAT), with bivariate analyses performed using t-tests.
Main Results:
- Gemini-generated materials showed improved readability (grade level 5.9) compared to baseline (7.7), while ChatGPT (12.5) and Copilot (8.8) performed worse.
- All LLM-generated PEMs achieved understandability scores above 70%, but performed worse than baseline materials (75%-83% vs. 75%-100%).
- Actionability scores for LLM-generated PEMs were significantly lower than baseline materials (40%-80% vs. 80%-100%).
Conclusions:
- There is significant variability in the performance of LLMs when generating de novo patient education materials (PEMs).
- While Gemini demonstrated enhanced readability and all LLMs met understandability targets, existing materials remain superior in both understandability and actionability.
- The integration of LLMs in creating PEMs should be carefully considered, balancing their potential benefits with essential clinical expertise to ensure optimal patient understanding and engagement.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Methods of Documentation II: POMR
Genetic Lingo
Patient-centered Care
Methods of Documentation III: PIE

