Related Experiment Video
Updated: Jan 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating Retrieval-Augmented Generation-Large Language Models for Infective Endocarditis Prophylaxis: Clinical
Paak Rewthamrongsris1, Vivat Thongchotchat2, Jirayu Burapacheep3
1Department of Anatomy, Faculty of Dentistry, Center of Artificial Intelligence and Innovation (CAII), Center of Excellence for Dental Stem Cell Biology, Chulalongkorn University, Bangkok, Thailand.
Retrieval-augmented generation (RAG) large language models (LLMs) show promise for infective endocarditis (IE) prophylaxis decision support. However, accuracy varies, and caution is advised for clinical use.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Medical Informatics
Background:
- Large language models (LLMs) are increasingly used in healthcare.
- Retrieval-augmented generation (RAG) enhances LLMs by grounding responses in specific, current data.
- Limitations of standard LLMs necessitate methods like RAG for reliable medical applications.
Purpose of the Study:
- To evaluate RAG-augmented LLMs for infective endocarditis (IE) prophylaxis recommendations in dental procedures.
- To compare the performance of RAG-LLMs against non-RAG LLMs using a standardized question set.
- To explore the utility of LLMs as a clinical decision support tool through a pilot study with dental students.
Main Methods:
- Ten RAG-integrated LLMs were tested using the 2021 American Heart Association IE guideline.
- A consistent IE prophylaxis question set from prior research was utilized for comparability.
- Performance was assessed with and without a preprompt, and a pilot study evaluated LLM assistance for dental students.
Main Results:
- Grok 3 beta achieved 90.0% accuracy with preprompting; DeepSeek Reasoner had the highest accuracy (83.6%) without preprompting.
- Preprompting generally improved LLM accuracy, though RAG's impact varied by model.
- The pilot study indicated mixed results for LLM assistance on accuracy and a significant increase in task time for students.
Conclusions:
- RAG and prompt engineering can improve LLM performance for clinical decision support.
- Current LLMs with RAG offer rapid information access but require critical evaluation due to potential inaccuracies.
- Clinicians and students must exercise digital literacy and maintain professional judgment when using these AI tools.
Related Concept Videos
Endocarditis III: Medical Management
Endocarditis I: Introduction
Endocarditis II: Clinical Features of Infective Endocarditis
Endocarditis IV: Nursing Management
Estimation of k and VD of Aminoglycosides
Healthcare Associated Infections II: Preventive Measures
The best practices for preventing healthcare-associated infections include hand hygiene, patient risk...

