Related Experiment Video
Updated: Oct 5, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Toward Responsible and Privacy-Preserving Generation of Inpatient Clinical Discharge Notes Using Large Language
Felice Burn1, Chantal Zwick2, Joseph Weibel2
1Center of Excellence for AI and Data Science, Cantonal Hospital Aarau, Aarau, Switzerland.
Abstract:
Writing discharge notes (DNs) is a time-intensive task in inpatient care that contributes substantially to physician workload and burnout. Large language models (LLMs) offer potential opportunities to automate clinical documentation. However, concerns regarding accuracy, reliability, data privacy protection laws and data governance, and compliance with healthcare privacy remain. This study evaluated the performance of locally deployed, domain-adapted LLMs for generating German-language DN in internal medicine. In this retrospective study, electronic health record data from a Swiss tertiary care hospital were used. Clinical information from 3 inpatient scenarios served as input for the 3 open source LLMs (Mixtral 8×7B, Mixtral 8×22B, and LLaMA 3.1 70B). Zero-shot prompting, 4-shot in-context learning, and supervised fine-tuning using QLoRA were compared. Generated DNs were evaluated using automated metrics, blinded physician assessment with direct comparison of all relevant clinical information using the modified Physician Documentation Quality Instrument, and additionally an LLM-as-a-judge framework and a Turing Test. Across all evaluation methods, the fine-tuned Mixtral 8×7B model achieved the best overall performance compared with prompting-based approaches and larger nonspecialized models. Simpler cases consistently received higher-quality ratings. Physician inter-rater agreement was low, highlighting the challenges of objective quality assessment for clinical documentation. In a Turing Test, physicians were unable to reliably distinguish artificial intelligence-generated from clinician-authored DNs, performing at chance level, demonstrating that high-quality DN generation is feasible using locally deployed LLMs within secure hospital infrastructures. Domain-specific fine-tuning substantially improves performance and may enable practical deployment. Future work should address further hallucination mitigation, robustness across medical specialties, and real-world integration.