Related Experiment Video
Updated: Jun 24, 2026

Eye-Tracking Control to Assess Cognitive Functions in Patients with Amyotrophic Lateral Sclerosis
Published on: October 13, 2016
Evaluating Retrieval Augmented Generation-enhanced Large Language Models for Question Answering On German
Marius Vach1, Michael Gliem2, Daniel Weiss3
1Department of Diagnostic and Interventional Radiology, Medical Faculty and University Hospital Düsseldorf, Heinrich-Heine-University Düsseldorf, Moorenstraße 5, 40225, Düsseldorf, Germany. Marius.Vach@med.uni-duesseldorf.de.
Retrieval-augmented Generation (RAG) significantly enhances Large Language Models (LLMs) for answering medical guideline questions. While RAG improves accuracy, further development is needed for safe clinical use, especially for non-English languages.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Medicine
- Clinical Guideline Application
Background:
- Large Language Models (LLMs) show promise for medical information retrieval.
- Accurate question answering based on complex medical guidelines remains a challenge.
- Retrieval-augmented Generation (RAG) offers a potential solution to enhance LLM performance.
Purpose of the Study:
- To evaluate the effectiveness of RAG-enhanced LLMs in answering questions related to German neurovascular guidelines.
- To compare the performance of different LLMs with and without RAG integration.
- To analyze the retrieval performance of various strategies for medical guideline information.
Main Methods:
- Four LLMs (GPT-4o-mini, Llama 3.1, Mixtral 8×22B, Claude 3.5 Sonnet) were tested with RAG.
- GPT-4o-mini without RAG served as a baseline.
- Answers were expert-validated, and retrieval strategies were assessed using a synthetic dataset.
Main Results:
- Claude 3.5 Sonnet demonstrated the highest accuracy (70.6%) among RAG-enhanced LLMs.
- RAG significantly improved LLM performance compared to models without RAG (20.0% accuracy).
- Retrieval errors constituted 80% of incorrect answers, with BM25 outperforming vector-based methods.
Conclusions:
- RAG substantially boosts LLM accuracy for medical guideline Q&A, though error rates persist.
- Further improvements in accuracy and confidence metrics are crucial for clinical deployment.
- General LLMs exhibit strong performance in non-English medical Q&A without specialized training.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Pre-Procedural Guidelines for Assessing Blood Pressure
Model Approaches for Pharmacokinetic Data: Physiological Models
Physiological Pharmacokinetic Models: Blood Flow-Limited Versus Diffusion-Limited Models
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...