Evaluating large language models for evidence-based clinical question answering

Can Wang1, Yiqun Chen2

  • 1Department of Biostatistics, Johns Hopkins University, Baltimore, MD 21205, USA.

Summary

Large language models (LLMs) show promise in medicine, but their reliability needs testing. Structured guidelines improve LLM accuracy, while retrieval augmentation enhances performance for evidence-based medicine.