Related Experiment Video
Updated: May 9, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Can a Large Language Model Grounded in Text-Based Agency-Specific Prehospital Protocols Provide Accurate Care
Colin G Wang1, Nichole Bosson1,2,3, Rombod Rahimian4
1Department of Emergency Medicine, Harbor-UCLA Medical Center, Torrance, California.
Retrieval-augmented generation (RAG) large language models (LLMs) achieved 75% accuracy in providing prehospital care recommendations based on emergency medical services (EMS) protocols. This study evaluated LLM accuracy for adult and pediatric emergency scenarios, finding no patient safety risks from identified hallucinations.
Area of Science:
- Artificial Intelligence in Healthcare
- Emergency Medical Services Research
- Clinical Decision Support Systems
Background:
- Large language models (LLMs) with retrieval-augmented generation (RAG) can provide source-grounded responses.
- Evaluating the accuracy of RAG-based LLMs for prehospital care recommendations is crucial for potential clinical integration.
Purpose of the Study:
- To assess the accuracy of a RAG-based LLM in generating prehospital care recommendations aligned with emergency medical services (EMS) policies and treatment protocols (TPs).
- To identify and categorize any inaccuracies or omissions in LLM-generated recommendations across diverse clinical scenarios.
Main Methods:
- A simulation-based study utilized Google's NotebookLM with a RAG-based LLM (Gemini 2.5 Flash).
- Text-based EMS policies/TPs from a large EMS system were uploaded.
- Six clinical scenarios (adult and pediatric) were developed, and LLM responses were independently evaluated for accuracy and hallucinations, categorizing missed actions by severity.
Main Results:
- The LLM provided 75% (127/169) of recommended patient care actions across all scenarios.
- 42 actions were missed, with 5% categorized as 'major misses' and 8% as 'minor misses'.
- Five major misses occurred in a pediatric out-of-hospital cardiac arrest scenario, primarily due to failure to prompt for secondary causes; 12 hallucinations were identified, none posing a safety risk.
Conclusions:
- A RAG-based LLM demonstrated 75% accuracy in generating prehospital care recommendations grounded in EMS policies and TPs.
- While generally accurate, specific areas like identifying secondary causes in pediatric emergencies require further refinement.
- The study highlights the potential of RAG-LLMs as decision support tools in prehospital care, with a need for continued validation and improvement.
Related Concept Videos
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic illness...
Types of Reports III: Telephone and Verbal Reports
Here's an overview of each type:
Telephone Orders
Cardiopulmonary Resuscitation II: ACLS Airway Management
Impact of Pharmacokinetic–Pharmacodynamic Models: Regulatory Decisions
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.
Cardiopulmonary Resuscitation IV: Pharmacological Management
