Related Experiment Video
Updated: May 26, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessing Completeness of Clinical Histories Accompanying Imaging Orders Using Adapted Open-Source and Closed-Source
David B Larson1,2, Arogya Koirala1,2, Lina Y Cheuy1,2
1Department of Radiology, Stanford University School of Medicine, 453 Quarry Rd, MC 5659, Stanford, CA 94304.
An adapted open-source large language model (LLM), Mistral-7B, effectively extracts clinical history elements from imaging orders, achieving performance comparable to GPT-4 Turbo. This tool establishes a benchmark for clinical history completeness, with 26.2% of histories containing all key elements.
Area of Science:
- Artificial Intelligence in Radiology
- Natural Language Processing for Clinical Data
Background:
- Incomplete clinical histories in radiology are a persistent issue, traditionally requiring manual analysis for quality improvement.
- Previous methods for assessing clinical history completeness are labor-intensive and not easily scalable.
Purpose of the Study:
- To adapt and evaluate open-source and closed-source large language models (LLMs) for automated extraction of clinical history elements from imaging orders.
- To utilize the best-performing open-source LLM to benchmark the completeness of a large dataset of clinical histories.
Main Methods:
- Retrospective analysis of 50,186 clinical histories from CT, MRI, US, and radiography orders.
- Adaptation of Llama 2-7B, Mistral-7B (open-source), and GPT-4 Turbo (closed-source) LLMs using prompt engineering, in-context learning, and fine-tuning.
- Evaluation of LLM performance against manual annotations by board-certified radiologists using accuracy, Cohen κ, and BERTScore.
Main Results:
- The fine-tuned open-source Mistral-7B model demonstrated substantial agreement with radiologists (κ = 0.73) and high semantic similarity (BERTScore = 0.96).
- Mistral-7B's performance closely rivaled that of GPT-4 Turbo in accuracy (91% vs. 92%) and semantic similarity.
- Using Mistral-7B, only 26.2% of clinical histories were found to contain all five essential elements (past medical history, what, when, where, clinical concern).
Conclusions:
- A fine-tuned open-source LLM (Mistral-7B) can effectively extract clinical history elements, matching the performance of larger closed-source models.
- This LLM provides a scalable solution for assessing clinical history completeness, establishing a crucial benchmark for clinical practice.
- The developed model and code will be fully open-sourced to promote wider adoption and further research.
More Related Videos
07:13Author Spotlight: An Efficient and Robust Software for Automated Fusion of Multiple Preclinical Imaging Modalities
Published on: October 27, 2023
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Related Concept Videos
Data Collection II
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body...