Related Experiment Video
Updated: Jan 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessing the Limitations of Large Language Models in Clinical Practice Guideline-Concordant Treatment
Tobias Roeschl1,2,3,4,5, Marie Hoffmann2,4,5, Djawid Hashemi1,2,3,4
1Department of Cardiology, Angiology and Intensive Care Medicine, Deutsches Herzzentrum der Charité, Berlin, Germany.
Large language models (LLMs) struggle with real-world clinical data for treatment decisions. Curated input is essential for accurate, unbiased LLM performance in healthcare.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
Background:
- Large language models (LLMs) show promise in therapeutic decision-making, but prior studies used highly curated patient data.
- The effectiveness of LLMs in real-world clinical scenarios with unstructured data remains largely uninvestigated.
Purpose of the Study:
- To evaluate the ability of various large language models (LLMs) to make guideline-concordant treatment decisions for severe aortic stenosis using typical clinical data.
- To assess LLM performance based on the input data format (unstructured medical reports vs. case summaries) and prompt augmentation.
Main Methods:
- A retrospective study of 80 severe aortic stenosis patients treated in 2022.
- Multiple LLMs were queried using anonymized original medical reports and manually generated case summaries.
- Performance was measured by agreement with institutional heart team decisions (Cohen κ), reliability (ICC), and fairness (FBI).
Main Results:
- LLMs demonstrated poor performance with original, unstructured medical reports (Cohen κ: -0.47 to 0.22).
- Performance significantly improved with case summaries and added guideline knowledge (Cohen κ: -0.02 to 0.63).
- All tested LLMs exhibited hallucinations, and bias towards TAVR was observed (FBI: 0.46-1.51).
Conclusions:
- Advanced LLMs require meticulously curated input for reliable clinical decision-making.
- The presence of hallucinations and bias necessitates caution when implementing LLMs in real-world healthcare settings.
- Further research is needed to enhance LLM accuracy and safety for clinical applications.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Related Concept Videos
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Language and Cognition
Analysis of Population Pharmacokinetic Data
Therapeutic Drug Monitoring: Affecting Factors