Related Experiment Video
Updated: Sep 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models for Clinical Note Simplification: A Systematic Review and Experimental Evaluation of Medical
Mara Teichmann1, Pelin Özkara Menekseoglu2, Julian Schwarz2
1Brandenburg University of Applied Science, Brandenburg a. d. Havel, Germany.
Abstract:
Since the introduction of the Patient Rights Act, patients in Germany have gained legal access to their medical records, including clinical notes. However, these documents are typically written for healthcare professionals and are often difficult for patients to understand due to specialized terminology, abbreviations, and complex sentence structures. Large Language Models (LLMs) offer new opportunities to automatically simplify such texts while preserving medically relevant information. This study investigates the potential of LLMs to improve the readability of German clinical notes by combining a systematic literature review with an experimental evaluation. Ten freely available LLMs were assessed using five synthetic clinical notes, which were simplified through standardized prompts designed to ensure linguistic clarity while maintaining content fidelity. Readability was analyzed using established indices (Flesch Reading Ease, Wiener Sachtextformel, LIX, SMOG, and Coleman-Liau) and complemented by a novel analysis of medical terminology and abbreviation density as indicators of domain-specific complexity. The results show that all LLMs substantially increased overall text length while consistently reducing the density of technical terms and abbreviations. However, no model achieved consistent improvements across all readability indices, highlighting limitations of traditional metrics in the medical domain. Models such as Mistral, ChatGPT, and Copilot demonstrated the highest efficiency in balancing linguistic simplification and text length. Overall, LLMs show strong potential to enhance the accessibility of clinical documentation for patients. However, their effectiveness depends on model selection, prompt design, and evaluation methodology. The findings suggest that conventional readability metrics should be extended with domain-specific measures to more accurately assess comprehensibility in medical texts.
