Related Experiment Video
Updated: Sep 16, 2025

The Dyspepsia Educational Tool As a Novel Aid in Dyspepsia Management
Published on: June 29, 2019
Improving the Readability of Institutional Heart Failure-Related Patient Education Materials Using GPT-4:
Ryan C King1, Jamil S Samaan2, Joseph Haquang1
1Department of Medicine, Division of Cardiology, University of California, Irvine Medical Center, 101 The City Dr S, Orange, CA, 92868, United States, 1 714-456-7890.
Insights
Large language models like GPT-4 can simplify heart failure patient education materials, improving readability and comprehension for better self-management. This AI tool enhances educational content accuracy and comprehensiveness for heart failure patients.
Area of Science:
- Medical Informatics
- Health Literacy
- Artificial Intelligence in Healthcare
Background:
- Heart failure management requires patient adherence to lifestyle changes, emphasizing the need for clear Patient Education Materials (PEMs).
- Existing cardiovascular PEMs often present information at a reading level exceeding recommended standards, hindering patient understanding.
- Large Language Models (LLMs) show potential for enhancing the accessibility of health information.
Purpose of the Study:
- To evaluate the readability of heart failure PEMs from leading cardiology institutions.
- To assess GPT-4's capability in improving PEM readability while maintaining accuracy and comprehensiveness.
Main Methods:
- Collected 143 heart failure PEMs from top US cardiology institutions.
- Utilized GPT-4 to simplify PEMs using a standardized prompt.
- Assessed readability using multiple metrics (Flesch Reading Ease, FKGL, etc.).
- Evaluated revised PEMs for accuracy and comprehensiveness by a cardiologist.
Main Results:
- GPT-4 significantly reduced the median Flesch-Kincaid Grade Level from 10.3 to 7.3 (p<.001).
- The proportion of PEMs below a sixth-grade reading level increased from 9.1% to 23.1% after GPT-4 revision (p<.001).
- No reduction in accuracy or comprehensiveness was observed; 23.1% of revised PEMs were deemed more comprehensive.
Conclusions:
- GPT-4 effectively enhances the readability of heart failure PEMs.
- LLMs may serve as valuable tools to supplement healthcare professional guidance for heart failure patients.
- Further research is essential to validate the safety, efficacy, and impact of AI in improving patient health literacy.
Background:
Heart failure management involves comprehensive lifestyle modifications such as daily weights, fluid and sodium restriction, and blood pressure monitoring, placing additional responsibility on patients and caregivers, with successful adherence often requiring extensive counseling and understandable patient education materials (PEMs). Prior research has shown PEMs related to cardiovascular disease often exceed the American Medical Association's fifth- to sixth-grade recommended reading level. The large language model (LLM) ChatGPT may be a useful tool for improving PEM readability.
Objective:
We aim to assess the readability of heart failure-related PEMs from prominent cardiology institutions and evaluate GPT-4's ability to improve these metrics while maintaining accuracy and comprehensiveness.
Methods:
A total of 143 heart failure-related PEMs were collected from the websites of the top 10 institutions listed on the 2022-2023 US News & World Report for "Best Hospitals for Cardiology, Heart & Vascular Surgery." PEMs were individually entered into GPT-4 (version updated July 20, 2023), preceded by the prompt, "Please explain the following in simpler terms." Readability was assessed using the Flesch Reading Ease score, Flesch-Kincaid Grade Level (FKGL), Gunning Fog Index, Coleman-Liau Index, Simple Measure of Gobbledygook Index, and Automated Readability Index. The accuracy and comprehensiveness of revised GPT-4 PEMs were assessed by a board-certified cardiologist.
Results:
For 143 institutional heart failure-related PEMs analyzed, the median FKGL was 10.3 (IQR 7.9-13.1; high school sophomore) compared to 7.3 (IQR 6.1-8.5; seventh grade) for GPT-4's revised PEMs (P<.001). Of the 143 institutional PEMs, there were 13 (9.1%) below the sixth-grade reading level, which improved to 33 (23.1%) after revision by GPT-4 (P<.001). No revised GPT-4 PEMs were graded as less accurate or less comprehensive compared to institutional PEMs. A total of 33 (23.1%) GPT-4 PEMs were graded as more comprehensive.
Conclusions:
GPT-4 significantly improved the readability of institutional heart failure-related PEMs. The model may be a promising adjunct resource in addition to care provided by a licensed health care professional for patients living with heart failure. Further rigorous testing and validation is needed to investigate its safety, efficacy, and impact on patient health literacy.
More Related Videos
Related Concept Videos
Heart Failure IV: Classification and Diagnostic Evaluation
Pathophysiology of Heart Failure
Heart Failure Drugs: β-Blockers
Heart Failure Drugs: Diuretics
Heart Failure VII: Nursing Interventions
Heart Failure Drugs: Inhibitors of Renin-Angiotensin System

