Related Experiment Video
Updated: Jun 3, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A dataset and benchmark for hospital course summarization with adapted large language models
Asad Aali1,2, Dave Van Veen3,4, Yamin Ishraq Arefeen2
1Department of Radiology, Stanford University, Stanford, CA 94304, United States.
Large language models (LLMs) can now synthesize brief hospital course (BHC) summaries from clinical notes. GPT-4 demonstrated superior performance in a clinical reader study, highlighting the potential of LLMs in healthcare summarization.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Informatics
- Natural Language Processing
Background:
- Brief hospital course (BHC) summaries are crucial clinical documents.
- Automating BHC synthesis from clinical notes using large language models (LLMs) is an emerging area.
- Existing LLM capabilities in healthcare summarization require further investigation.
Purpose of the Study:
- To introduce the MIMIC-IV-BHC dataset for adapting LLMs to BHC synthesis.
- To benchmark the summarization performance of general-purpose and healthcare-adapted LLMs.
- To evaluate LLM-generated BHCs for clinical decision-making enhancement.
Main Methods:
- Utilized clinical notes as input for LLMs.
- Applied prompting-based and fine-tuning-based adaptation strategies.
- Evaluated LLMs using quantitative metrics (BLEU, BERT-Score) and a qualitative clinical reader study with 5 clinicians.
Main Results:
- Fine-tuned Llama2-13B outperformed other domain-adapted models on quantitative metrics.
- GPT-4 with in-context learning showed robustness to increasing context lengths.
- Clinicians significantly preferred GPT-4 generated summaries over fine-tuned Llama2-13B and original summaries (P<.001).
Conclusions:
- The MIMIC-IV-BHC dataset and LLM performance benchmark are released.
- Both proprietary and open-source LLMs show high-quality summarization performance.
- LLM-generated summaries, particularly from GPT-4, show potential for improving clinical decision-making.
Related Concept Videos
Hospitals-II
Nurses that work in...
Improving Translational Accuracy
Hospitals-I
Statistical Software for Data Analysis and Clinical Trials
Kaplan-Meier Approach
Truncation in Survival Analysis
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...

