Related Experiment Video
Updated: Mar 20, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Automated auditing of emergency department documentation using large language models
Nicholas Dietrich1, Brett Stubbert2
1Temerty Faculty of Medicine, University of Toronto, 1 King's College Cir, Toronto, ON M5S 1A8, Canada.
Introduction:
Errors in emergency department (ED) documentation can lead to patient harm and medicolegal risk, however manual document auditing is resource-intensive and difficult to scale. Large language models (LLMs) may offer an automated alternative, but their performance for task-specific error detection remains uncertain. This study evaluated four state-of-the-art LLMs for automated auditing of ED documentation.
Methods:
A total of 600 ED discharge notes from real-world encounters were evaluated, including 500 notes with a single injected documentation error across five predefined categories (patient/encounter, medication/allergy, clinical findings, investigation failure/contradiction, or disposition/follow-up), and 100 error-free reference notes. Four LLMs (Claude Sonnet 4.5, Gemini 2.5 Pro, GPT-5, and Grok 4) independently audited each note under blinded conditions. Model performance was assessed using accuracy, precision, recall, and F1 score, with pairwise comparisons performed using McNemar's test.
Results:
Model accuracies ranged from 93.2% to 94.2% for detection of any documentation error, with precision exceeding 98% and F1 scores above 0.95 for all models. Claude Sonnet 4.5 achieved the highest overall accuracy (94.2%) and F1 score (0.964). Pairwise comparisons showed no statistically significant differences in overall error detection between models (p > 0.05). At the category level, all models demonstrated near-perfect performance for patient/encounter inconsistencies and medication/allergy errors (F1 scores >0.96), with lower performance for investigation failures/contradictions, followed by disposition/follow-up errors and clinical finding inconsistencies.
Conclusion:
LLMs demonstrate strong performance in auditing ED documentation, supporting their potential role as quality assurance tools in emergency care. Future work should explore how these systems can be integrated into real clinical workflows.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Related Concept Videos
Methods of Documentation VII: EMR
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.
Methods of Documentation V: CBE
In CBE, healthcare professionals establish predefined standards of practice that define what constitutes...
Methods of Documentation II: POMR
Documentation in Long-Term and Home Healthcare Setting
Long-Term Care Facilities