A Self-Controlled Benchmark of Retrieval-Augmented Generation for Large Language Models on Clinical Guideline

Andreas Vollmer1, Lara Schorn2, Felix Schrader2

  • 1Department of Oral and Maxillofacial Plastic Surgery, University Hospital of Wuerzburg, 97070 Wuerzburg, Germany.

Summary

Retrieval-augmented generation (RAG) significantly improves large language model (LLM) accuracy and safety for clinical guideline questions. This enhancement is reproducible and base-independent, especially for weaker models, reducing hallucinations.

Related Concept Videos