Related Experiment Video
Updated: Jun 23, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparison of Large Language Model-Based Systems and Prompt Engineering for Internal Medicine Clinical Pharmacy Cases
Samuel S Yang1, Clement E Ng1, Hyunuk Seung2
1University of Maryland Medical Center, Baltimore, Maryland, USA.
Evaluating artificial intelligence in clinical pharmacy, this study found neither domain-specific large language models (LLMs) nor prompt engineering improved response accuracy or completeness. However, one LLM demonstrated superior reference validity for practice applications.
Area of Science:
- Artificial Intelligence in Healthcare
- Clinical Pharmacy Informatics
- Large Language Model (LLM) Applications
Background:
- Large Language Models (LLMs) are AI tools generating responses with variable accuracy and completeness.
- Methods to enhance LLM response quality in clinical pharmacy practice require further evaluation.
Purpose of the Study:
- To evaluate the impact of domain-specific LLMs versus general-purpose LLMs on response quality.
- To assess the utility of prompt engineering templates in refining LLM outputs for clinical pharmacy.
- To compare the accuracy, completeness, and reference validity of different LLM systems.
Main Methods:
- A prospective, observational study utilized 50 clinical case questions.
- Two LLM systems were compared: ChatGPT 4o (general-purpose) and OpenEvidence (healthcare-specific with RAG).
- A two-by-two factorial design assessed the effect of prompt engineering templates on LLM responses, evaluated by pharmacists.
Main Results:
- No statistically significant interaction was found between LLM type and prompt engineering for accuracy or completeness.
- OpenEvidence demonstrated significantly higher reference validity compared to ChatGPT (p < 0.001).
- Predicted probabilities for the primary outcome (accuracy and completeness) varied across conditions, with OpenEvidence without a template showing 0.64.
Conclusions:
- Neither domain-specific LLMs with RAG nor prompt engineering templates significantly improved LLM accuracy or completeness in this clinical pharmacy context.
- OpenEvidence's superior reference validity suggests potential for clinical practice integration.
- Further research is needed to optimize LLM utility as AI adoption grows in healthcare.
Related Concept Videos
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Introduction to Language of Pathophysiology ll
Pharmacodynamic Models: Link Model and Systems Pharmacodynamic Model
Pharmacodynamic Models: Overview
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic illness...
Pharmacodynamic Models: Direct Effect Model and Indirect Response Model