Related Experiment Video
Updated: Apr 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating the Performance of Large Language Models for Generating Emergency Department Discharge Instructions
Patricia Hernández1, Giovanni Rodriguez2, Chanel Fischetti2
1Department of Emergency Medicine, Massachusetts General Hospital, Brigham and Women's Hospital, Harvard Medical School, Boston, Massachusetts.
Background:
Clear discharge instructions are essential for safe emergency department (ED) care, yet creating patient-centered, health literate materials remain challenging. Large language models (LLMs) may improve patient communication, but their role in the ED remains underexplored.
Objective:
The objective of this study was to evaluate the feasibility of using LLMs to generate ED discharge instructions and to compare readability across models.
Methods:
We generated discharge instructions for 20 ED diagnoses across Emergency Severity Index (ESI) levels using GPT-4, GPT-4o, GPT-5.2, Gemini 2.5 Pro, Gemini 3 Pro, Gemini 3 Flash, and compared them with available electronic medical record (EMR) stock discharge templates. Readability was assessed with eight indices, and differences across ESI levels were tested via Kruskal-Wallis analysis. Diagnosis-level paired comparisons versus stock discharges were performed with effect sizes and multiplicity adjustment. Four blinded emergency medicine attendings evaluated GPT-4 generated instructions for medical accuracy, clarity, completeness, and understandability using a validated survey. Inter-rater reliability (IRR) was calculated using the kappa statistic.
Results:
Across models, readability frequently exceeded recommended 6th-8th grade health-literacy targets despite prompting. Compared with stock templates, readability differences varied by model and index, with Gemini 2.5 Pro and Gemini 3 Pro generally producing the lowest grade-level outputs, whereas GPT-5.2 produced the highest. We observed a general trend that instructions for lower ESI levels had poorer readability. Physician ratings generally exceeded 4/5 in most domains, with > 50% rated 5/5 for clarity, understandability, completeness, and accuracy. However, approximately 25% of instructions scored ≤ 3/5 for completeness or understandability. IRR was fair (κ = 0.40), with lower agreement for lower ESI levels.
Conclusions:
LLMs can generate ED discharge instructions with strong clinician-rated quality in many domains, but readability performance varies substantially by model and metric and often fails to meet health-literacy standards. Careful model selection, prompt refinement, and human oversight remain necessary to ensure accessible and complete discharge communication. With further refinements, LLMs could improve patient comprehension and support safe discharge practices in the ED.
Related Concept Videos
Discharge Summary Forms
Here's a detailed look at the key components and guidelines for preparing a discharge summary:
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.
