Related Experiment Video
Updated: Jun 3, 2026

07:31
Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Open-source large language model-based on-premises pipeline for automated data extraction from unstructured
Vasileios Ntinopoulos1,2, Hector Rodriguez Cetina Biefer1,2,3, Laura Rings1,2
1Department of Cardiac Surgery, University Hospital Zurich, Zurich, Switzerland.
BMJ Health & Care Informatics
|June 1, 2026
Summary
Open-source large language models (LLMs) show high accuracy for automated data extraction from electronic health records (EHRs). This pilot study confirms the feasibility of privacy-preserving LLM pipelines for healthcare applications.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Natural Language Processing
Background:
- Unstructured electronic health records (EHRs) contain valuable patient data.
- Automated data extraction from EHRs is crucial for clinical research and improved patient care.
- Current methods for EHR data extraction are often manual and time-consuming.
Purpose of the Study:
- To evaluate an on-premises, open-source large language model (LLM)-based pipeline for automated data extraction from unstructured EHRs.
- To assess the performance of various LLMs in different data extraction and classification tasks.
- To determine the reliability and consistency of LLM responses for EHR data.
Main Methods:
- Preprocessed 50 German medical texts using an automated script.
- Evaluated 14 mid-sized LLMs (30B-90B parameters) with 4-bit and 8-bit quantization.
- Assessed LLM performance across information extraction, binary classification, and multilevel classification tasks for the EuroSCORE 2 model.
- Measured LLM response consistency over three iterations.
Main Results:
- Multiple LLMs achieved high accuracy (over 0.90) in overall and classification tasks.
- 12 LLMs demonstrated perfect (1.0) accuracy in information extraction.
- Nine LLMs showed perfect response consistency, with others achieving Krippendorff's alpha of 0.999.
- Qwen3-30b-a3b-q8 and Llama3.2-vision-90b-q4 showed top performance in specific tasks.
Conclusions:
- LLMs can reliably automate data extraction from EHRs with high accuracy.
- On-premises, privacy-preserving LLM pipelines are feasible for healthcare.
- Further large-scale studies are needed to validate LLM applications in clinical settings.

