Related Experiment Video
Updated: Jun 18, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
534
Evaluating local open-source large language models for data extraction from unstructured reports on mechanical
Aymen Meddeb1,2, Philipe Ebert3, Keno Kyrill Bressem4
1Department of Neuroradiology, Charité Universitätsmedizin Berlin, Berlin, Germany aymen.meddeb@charite.de.
Journal of Neurointerventional Surgery
|August 2, 2024
Summary
Open-source large language models (LLMs) show promise for extracting clinical data from mechanical thrombectomy reports. Combining LLMs with human-in-the-loop (HITL) improves accuracy and saves significant time in data extraction for stroke research.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Clinical Research
Background:
- Assessing the efficacy of open-source large language models (LLMs) for clinical data extraction from mechanical thrombectomy reports.
- Focusing on patients with ischemic stroke due to vessel occlusion.
Purpose of the Study:
- To evaluate the performance of local open-source LLMs in extracting key clinical data points.
- To compare the efficiency of automated versus manual data extraction using a human-in-the-loop (HITL) approach.
Main Methods:
- Deployed three LLMs (Mixtral, Qwen, BioMistral) on institutional and external datasets of mechanical thrombectomy reports.
- Utilized HITL for ground truth labeling and recorded time metrics for data extraction.
- Assessed model performance based on precision, recall, and F1 score across 15 clinical categories.
Main Results:
- Mixtral demonstrated high precision (0.99 for first series time).
- HITL approach resulted in an average time saving of 65.6% per case.
- Performance varied across models and data categories, with NIHSS scores and occluded vessels showing diverse extraction accuracy.
Conclusions:
- LLMs offer significant potential for automated clinical data extraction from unstructured medical reports.
- HITL integration enhances data precision and reliability, supporting clinical documentation and research.
- This privacy-preserving methodology is a scalable solution for healthcare data management.

