Related Experiment Video
Updated: Aug 14, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A Self-Controlled Benchmark of Retrieval-Augmented Generation for Large Language Models on Clinical Guideline
Andreas Vollmer1, Lara Schorn2, Felix Schrader2
1Department of Oral and Maxillofacial Plastic Surgery, University Hospital of Wuerzburg, 97070 Wuerzburg, Germany.
Diagnostics (Basel, Switzerland)
|August 13, 2026
Summary
Retrieval-augmented generation (RAG) significantly improves large language model (LLM) accuracy and safety for clinical guideline questions. This enhancement is reproducible and base-independent, especially for weaker models, reducing hallucinations.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
Background:
- Large language models (LLMs) show potential for clinical decision support but struggle with specialized medical guidelines.
- Retrieval-augmented generation (RAG) may improve LLM performance by grounding responses in authoritative knowledge bases.
Purpose of the Study:
- To compare the accuracy, comprehensiveness, and safety of RAG-enhanced versus standard LLMs for answering clinical questions from the German S3 guideline for oral cavity carcinoma.
- To evaluate the impact of retrieval-augmented generation on LLM performance in a specialized medical domain.
Main Methods:
- A prospective, single-blind benchmark study evaluated six LLMs (one RAG-enhanced, one consensus-based, four standard).
- Fifty clinical questions were posed to each model, with responses assessed by expert reviewers and an automated judge.
- A paired within-model experiment compared LLM performance with and without guideline retrieval via a transparent pipeline.
Main Results:
- Guideline retrieval significantly improved accuracy in weaker open-weight models and directionally in stronger ones.
- Content-level hallucination decreased from 42% to 4%, and accuracy rose by +0.64 points on the automated judge.
- Citation groundedness increased from 0% to 51-89%, with retrieval recall@5 at 92%.
Conclusions:
- Guideline retrieval offers a reproducible, base-independent improvement in LLM safety and auditability for clinical guideline questions.
- Accuracy gains are most pronounced in weaker base models, while hallucination and auditability improvements are consistent across models.
- Rater-independent evaluation measures are crucial due to expert recognizability of RAG-enhanced answers, and human oversight remains necessary for residual hallucinations.