Related Experiment Video
Updated: Jan 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating Large Language Model Performance in Generating Clinically Relevant Intensive Care Unit Discharge Summaries
Seshadri C Mudumbai1,2, Philip Chung2, Ji-Qing Chen1
1From the Anesthesiology, Perioperative and Pain Medicine Service, Veterans Affairs Palo Alto Health Care System, Palo Alto, California.
Background:
Large language models (LLMs) have the potential to automate time consuming clinical documentation tasks. One such task is generating ICU discharge summaries, which requires summarizing an often, complex clinical course and is time consuming. We compared the quality of LLM-generated ICU discharge summaries using expert intensivist evaluations. LLM summaries were assessed across six domains (coherence, consistency, fluency, relevance, utility, and overall quality relative to human-authored summaries) to determine their clinical suitability.
Methods:
Ten patient cases were randomly selected from the MIMIC-III database. Each case included exactly 20 physician notes (progress, consultation, and admission notes). The Bidirectional and Auto-Regressive Transformer (BART) model generated individual note summaries that were then combined into comprehensive LLM-written discharge summaries. Four experienced intensivists scored LLM-generated summaries on a 5-point Likert scale across the six domains.
Results:
LLM-generated summaries achieved median (IQR) scores of 4 (3-5) for coherence, 4 (3-5) for fluency, 3 (3-4) for consistency, 3 (2-4) for relevance, 3 (3-4) for utility, and 2 (2-3) for overall quality relative to human-authored summaries. Inter‑rater reliability was moderate for coherence (ICC = 0.62), consistency (0.65), and fluency (0.68), but lower for relevance (0.45) and utility (0.48). Although current LLMs were reasonably coherent, they frequently omitted patient‑specific information and scored lowest on utility, relevance and quality relative to human authored summaries.
Conclusions:
LLMs can generate fluent and readable ICU discharge summaries but may overlook critical clinical details and lack depth compared to human-authored summaries. Further ICU-specific fine-tuning and incorporation of domain-specific knowledge are needed to improve LLM alignment with human expertise.
Related Concept Videos
Discharge Summary Forms
Here's a detailed look at the key components and guidelines for preparing a discharge summary:
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.
Documentation in Long-Term and Home Healthcare Setting
Long-Term Care Facilities
