Related Experiment Video
Updated: May 26, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Validating large language model-assisted data extraction from clinical notes
J W van Koevorden1,2, N Aben3, V Struben3
1Department of Head and Neck Surgery, Antoni van Leeuwenhoek-Netherlands Cancer Institute, Amsterdam, the Netherlands.
Summary
Large language models (LLMs) significantly reduce clinical documentation time in head and neck oncology, achieving high accuracy. Human oversight is crucial for AI-assisted documentation tools to support clinical workflows effectively.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Informatics
- Natural Language Processing
Background:
- Healthcare professionals face significant documentation burdens impacting efficiency and patient safety.
- Large language models (LLMs) offer a potential scalable solution for automating data extraction from unstructured clinical notes.
- This study assesses LLM accuracy and clinical impact for structured data extraction in head and neck oncology.
Purpose of the Study:
- To compare the accuracy and efficiency of LLM-driven data extraction against manual physician extraction.
- To evaluate the clinical impact and error types associated with LLM data extraction.
- To determine the potential of LLMs in streamlining clinical documentation for oncology consultations.
Main Methods:
- A prospective validation study analyzed 1482 pages of clinical documentation from 60 patients.
- A pretrained open-source LLM and two physicians extracted data across 29 categories.
- Six clinical experts evaluated 2555 extracted values for accuracy, precision, recall, and F1 scores, categorizing errors.
Main Results:
- LLM extraction accuracy ranged from 74% (pathology) to 90% (patient characteristics).
- Manual extraction exhibited 29% interobserver disagreement, while LLM hallucinations were rare (0.16%) and low impact.
- LLM extraction reduced average case time from 8.6 to 1.9 minutes (P < 0.001).
Conclusions:
- LLMs can effectively reduce documentation time in clinical settings while maintaining acceptable accuracy.
- Human oversight is essential when implementing AI-assisted documentation tools.
- Findings support further research into AI tools for clinical practice to enhance documentation efficiency.
