Benchmarking large language models for question answering on German clinical practice guidelines

Johannes Schwietering1, Gregor Lichtner2

  • 1UMIT TIROL - Private University for Health Sciences and Health Technology, Hall in Tirol, Tyrol, Austria.

Summary

A new benchmark, cpgQA-DE, was created to evaluate German clinical practice guideline question answering. Retrieval-Augmented Generation (RAG) significantly improved large language model (LLM) accuracy, showing promise for clinical decision support.

Related Concept Videos