Related Experiment Video
Updated: Apr 29, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Are LLM-generated plain language summaries truly understandable? A large-scale crowdsourced evaluation.
Yue Guo1, Jae Ho Sohn2, Gondy Leroy3
1School of Information Sciences, University of Illinois Urbana-Champaign, Champaign, IL, USA.
Large language models (LLMs) can generate plain language summaries (PLSs) that seem clear but do not improve patient comprehension as effectively as human-written summaries. Automated metrics also fail to capture true understanding, highlighting the need for better AI evaluation in health communication.
Area of Science:
- Health Communication
- Artificial Intelligence in Medicine
- Patient Education
Background:
- Plain language summaries (PLSs) are crucial for patient understanding of medical information.
- Large language models (LLMs) show potential for automating PLS generation.
- Previous evaluations of LLM-generated PLSs lack robust measures of comprehension and generalizability.
Purpose of the Study:
- To conduct a large-scale crowdsourced evaluation of LLM-generated PLSs.
- To assess both subjective quality and objective reader performance (comprehension and recall).
- To examine the alignment between automated metrics and human judgments of PLS quality.
Main Methods:
- A crowdsourced study involving 150 participants via Amazon Mechanical Turk.
- Assessment of PLS quality via perceived ratings (simplicity, informativeness, coherence, faithfulness).
- Task-based measures including multiple-choice accuracy for comprehension and recall tests.
Main Results:
- Participants rated LLM-generated PLSs similarly to human-written ones in clarity and coherence.
- However, participants demonstrated significantly better comprehension after reading human-written PLSs.
- Automated evaluation metrics showed poor alignment with human judgments of PLS quality.
Conclusions:
- LLM-generated PLSs may appear fluent and trustworthy but do not reliably enhance patient comprehension.
- Current automated metrics are inadequate for evaluating the true effectiveness of PLSs.
- Future research must focus on developing generation methods and evaluation frameworks that prioritize layperson comprehension.
Related Concept Videos
Sources of Law
Constitutional law is foundational, deriving from federal and state constitutions, and...
Improving Translational Accuracy
Guidelines for Writing Outcome
Patient outcomes reflect the patient's response to the goal rather than what the nurse aims to achieve. Terminology should be observable and measurable to avoid the reader's interpretation. The desired outcome should be realistic and achievable in the designated care timeframe. Expected outcomes should align with adjunctive therapies. The outcome should enhance care...
Ligand Binding Sites
Health Literacy
Regulation of Expression Occurs at Multiple Steps
