Related Experiment Video
Updated: Jun 30, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Optimising the clinical application of rheumatology guidelines using large language models: a retrieval-augmented
Alfredo Madrid-García1, Diego Benavent2, Chamaida Plasencia-Rodríguez3
1Independent Researcher.
Objectives:
Timely access to current rheumatology guidelines at the point of care is challenging. We aimed to develop and evaluate the first retrieval-augmented generation (RAG) system designed for adult rheumatology, integrating European Alliance of Associations for Rheumatology (EULAR) and American College of Rheumatology (ACR) guidelines to provide rheumatologists with timely evidence-based recommendations.
Methods:
EULAR and ACR management guidelines were selected by rheumatologists based on their clinical relevance for decision-making and processed. A RAG system was implemented. To evaluate it, 10 questions per guideline were generated using ChatGPT 4.5. Answers to these were produced by ChatGPT-o3-mini with context retrieval (RAG) and without (baseline). Performance was assessed by an Large language model (LLM)-as-a-judge (Gemini 2.0 Flash) using a 5-point Likert scale across 5 dimensions: relevance, factual accuracy, safety, completeness, and conciseness; it also determined preference between the RAG and baseline responses. For validation, 2 rheumatologists independently evaluated a random sample of questions (15%) on the same domains. Statistical significance was established using the Wilcoxon signed-rank and binomial tests.
Results:
Seventy-four guidelines were included, yielding 740 evaluation questions. The LLM-as-a-judge evaluation showed the RAG system significantly outperformed the baseline across all criteria (P < .001). Manual evaluation confirmed these findings (P < .001) for accuracy, safety, and completeness. The RAG system was significantly preferred by the LLM-as-a-judge in 92.8% of comparisons and by the human evaluators in 71.2% to 74.8% (P < .001).
Conclusions:
Developing a RAG system integrating extensive EULAR/ACR rheumatology guidelines improves answer quality compared to a baseline LLM. This evaluation provides a robust foundation for reliable, artificial intelligence-driven clinical decision support tools designed to enhance evidence-based practice by providing clinicians with rapid, context-aware access to recommendations.
Related Concept Videos
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Rheumatic Heart Disease III: Medical Management
Improving Translational Accuracy
Improving Translational Accuracy
