Related Experiment Video
Updated: Mar 20, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Automated auditing of emergency department documentation using large language models
Nicholas Dietrich1, Brett Stubbert2
1Temerty Faculty of Medicine, University of Toronto, 1 King's College Cir, Toronto, ON M5S 1A8, Canada.
Large language models (LLMs) show high accuracy in auditing emergency department (ED) documentation errors. These AI tools can enhance quality assurance in emergency care by reliably detecting issues in clinical notes.
Area of Science:
- Health Informatics
- Artificial Intelligence in Healthcare
- Clinical Documentation Improvement
Background:
- Manual auditing of emergency department (ED) documentation is resource-intensive and difficult to scale.
- Errors in ED documentation pose risks to patient safety and increase medicolegal liability.
- The efficacy of large language models (LLMs) for specialized documentation error detection is not well-established.
Purpose of the Study:
- To evaluate the performance of four state-of-the-art large language models (LLMs) in the automated auditing of emergency department (ED) discharge notes.
- To assess the accuracy, precision, recall, and F1 scores of LLMs in identifying predefined documentation errors.
Main Methods:
- 600 ED discharge notes were analyzed, comprising 500 notes with single injected errors and 100 error-free controls.
- Five error categories were defined: patient/encounter, medication/allergy, clinical findings, investigation failure/contradiction, and disposition/follow-up.
- Four LLMs (Claude Sonnet 4.5, Gemini 2.5 Pro, GPT-5, Grok 4) performed blinded audits; performance metrics were calculated, and McNemar's test was used for comparisons.
Main Results:
- All evaluated LLMs achieved high accuracy (93.2%-94.2%) and F1 scores (>0.95) for overall error detection.
- Claude Sonnet 4.5 demonstrated the highest accuracy (94.2%) and F1 score (0.964); no significant differences were found between models (p > 0.05).
- Models excelled in detecting patient/encounter and medication/allergy errors (F1 > 0.96) but showed lower performance for investigation, disposition, and clinical finding errors.
Conclusions:
- Large language models show significant potential as automated quality assurance tools for emergency department documentation.
- The findings support the integration of LLMs into clinical workflows to improve the accuracy and safety of emergency care documentation.
- Further research is needed to optimize LLM performance across all error categories and facilitate seamless clinical workflow integration.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Related Concept Videos
Methods of Documentation VII: EMR
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.
Methods of Documentation V: CBE
In CBE, healthcare professionals establish predefined standards of practice that define what constitutes...
Methods of Documentation II: POMR
Documentation in Long-Term and Home Healthcare Setting
Long-Term Care Facilities