Related Experiment Video
Updated: Sep 19, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating large language models for drafting emergency department encounter summaries.
Christopher Y K Williams1, Jaskaran Bains2, Tianyu Tang2
1Bakar Computational Health Sciences Institute, University of California, San Francisco, California, United States of America.
Large language models (LLMs) can summarize clinical notes, but errors like hallucinations and omissions occur. While most errors have low harm potential, careful clinician review is essential for patient safety.
Area of Science:
- Clinical Informatics
- Artificial Intelligence in Healthcare
- Medical Documentation
Background:
- Large language models (LLMs) show promise for clinical applications like text summarization.
- The increasing deployment of AI scribes necessitates rigorous evaluation of their accuracy in healthcare settings.
Purpose of the Study:
- To evaluate the performance of GPT-4 and GPT-3.5-turbo in generating Emergency Department (ED) encounter summaries.
- To identify the prevalence and types of errors (inaccuracy, hallucination, omission) in LLM-generated ED summaries.
Main Methods:
- Cross-sectional study of 100 randomly sampled adult ED visits (2012-2023).
- Evaluation of GPT-4 and GPT-3.5-turbo generated summaries against three criteria: inaccuracy, hallucination, and omission.
- Analysis of error types and locations within encounter summary sections.
Main Results:
- GPT-4 generated error-free summaries in 33% of cases; GPT-3.5-turbo in 10%.
- GPT-4 summaries had 10% inaccuracies, 42% hallucinations, and 47% omissions.
- Inaccuracies/hallucinations were common in 'Plan' sections; omissions in 'Physical Examination' and 'History of Presenting Complaint' sections.
- Mean potential harm score for errors was low (0.57/7).
Conclusions:
- LLMs can generate clinical encounter summaries, but are prone to hallucinations and omissions.
- Errors in LLM-generated clinical text, while often low-harm, require careful clinician review.
- Understanding error patterns is crucial for safe integration of LLMs in clinical workflows.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Related Concept Videos
Methods of Documentation VII: EMR
Types of Reports III: Telephone and Verbal Reports
Here's an overview of each type:
Telephone Orders
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.
SBAR II: Application of SBAR
SBAR Report from a Nurse to a Health Care Provider
S: "Hello, Dr. Smith. This is Jane, RN, from the Med Surg unit. I am calling to tell you about Ms. White in Room 210, who is experiencing increased pain and redness at her incision site. Her recent...
Methods of Documentation V: CBE
In CBE, healthcare professionals establish predefined standards of practice that define what constitutes...