Related Experiment Video
Updated: May 24, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Advancing Knowledge in Evaluating the Clinical Impact of Large Language Models for Clinical Text Summarization: A
Lydie Bednarczyk1,2, Mina Bjelogrlic1,2, Jamil Zaghir1,2
1Division of Medical Information Sciences, Geneva University Hospitals, Geneva, Switzerland.
Large language models (LLMs) show promise for summarizing electronic health records (EHRs), but errors pose patient safety risks. Current evaluations of LLM summarization impact are inconsistent and lack standardization for safe clinical use.
Area of Science:
- Clinical Informatics
- Artificial Intelligence in Healthcare
- Medical Documentation
Background:
- Large language models (LLMs) offer potential to streamline clinical text summarization from electronic health records (EHRs), reducing physician documentation burden.
- However, inaccuracies in LLM-generated summaries can lead to misinformation, misrepresentation of clinical facts, and potential patient safety risks.
Purpose of the Study:
- To systematically review and analyze the evaluation methodologies of clinical impact for LLM-based summarization systems.
- To assess how studies have examined utility, failure modes, patient safety risks, and bias in LLM-generated clinical summaries.
Main Methods:
- A literature search was conducted in PubMed for studies published between June 2024 and September 2025.
- A total of 144 studies were retrieved, with 32 selected for in-depth analysis based on predefined criteria.
- Evaluations were categorized across four dimensions: utility, failure, patient safety risks, and bias.
Main Results:
- Failure analysis (50%) was most common, focusing on inaccuracies, omissions, and hallucinations with varied definitions and methods.
- Utility analysis (25%) assessed workflow efficiency, readability, and understandability.
- Bias analysis (22%) examined gender, race, and stigmatizing language; patient safety risk analysis was minimal (3%).
Conclusions:
- Current evaluation frameworks for LLM summarization are evolving beyond technical metrics but lack standardization.
- Heterogeneous methodologies hinder consistent assessment of clinical impact and safe deployment.
- Future research requires standardized error taxonomies, risk assessment frameworks, and bias detection methods for reliable clinical integration.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Introduction to Language of Pathophysiology ll
Clinical Trials: Overview