Related Experiment Video
Updated: Jun 3, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Open-source large language model-based on-premises pipeline for automated data extraction from unstructured
Vasileios Ntinopoulos1,2, Hector Rodriguez Cetina Biefer1,2,3, Laura Rings1,2
1Department of Cardiac Surgery, University Hospital Zurich, Zurich, Switzerland.
Objectives:
We evaluated an on-premises, open-source large language model (LLM)-based data extraction pipeline for automated data extraction from unstructured electronic health records (EHRs).
Methods:
Automated script-based EHR data preprocessing extracted 50 medical texts in German, which were entered into the LLM pipeline. 4-bit and 8-bit quantizations of 14 mid-sized LLMs (30B-90B parameters) were evaluated in 6 information extraction, 11 binary classification and 5 multilevel classification tasks comprising all variables of the European System for Cardiac Operative Risk Evaluation 2 (EuroSCORE 2) model for 1100 predictions each. LLM response consistency was assessed over three same-prompt iterations.
Results:
In overall accuracy, Qwen3-30b-a3b-q8 presented the highest value (0.954) and 13 LLMs had values over 0.90. In information extraction accuracy, 12 LLMs exhibited a value of 1.0 and all 14 LLMs had values over 0.96. In binary classification accuracy, Llama3.2-vision-90b-q4 exhibited the highest value (0.972), 5 LLMs had values of at least 0.95 and all 14 LLMs showed values over 0.93. In multilevel classification accuracy, Qwen3-30b-a3b-q8 exhibited the highest value (0.940), four LLMs had values over 0.90 and all LLMs presented values over 0.80. Nine LLMs exhibited perfect response consistency and the remaining five LLMs had a Krippendorff's alpha value of 0.999.
Discussion:
Multiple LLMs exhibited high accuracy in information extraction, binary classification, multilevel classification and response consistency and seem able to reliably automate data extraction from EHRs.
Conclusion:
This pilot study demonstrates the feasibility of on-premises, privacy-preserving, LLM-based automated EHR data extraction pipelines. Larger-scope studies are warranted to validate their potential in healthcare.

