Related Experiment Video
Updated: Apr 30, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.3K
Rise of the Machines: Comparing Performance of Artificial Intelligence Large Language Models on Pharmacy Specialty
Curtis D Collins1, Michael P Veve2,3, Ian B Hollis4
1Department of Pharmacy Services, Trinity Health Ann Arbor, Ann Arbor, Michigan, USA.
Journal of the American College of Clinical Pharmacy : JACCP
|April 29, 2026
Summary
Large language models (LLMs) show high accuracy on pharmacy certification exams, with top performers like Microsoft Copilot (GPT-5) excelling. Continued evaluation is recommended for clinical decision support.
Area of Science:
- Artificial Intelligence in Healthcare
- Clinical Decision Support Systems
- Pharmacy Practice
Background:
- Large language models (LLMs) are increasingly utilized for clinical information retrieval and decision support.
- Comparative performance of LLMs on specialized pharmacy content is not well-understood.
Purpose of the Study:
- To evaluate the performance of 15 large language models (LLMs) on Board of Pharmacy Specialties (BPS) certification practice questions.
- To identify differences in LLM accuracy across 14 pharmacy specialty domains.
Main Methods:
- 145 publicly available BPS certification practice questions across 14 specialties were used.
- 15 LLMs were tested using a standardized prompt without prompt engineering.
- Model responses were scored against official answer keys, with statistical analysis performed on accuracy differences.
Main Results:
- Mean accuracy across all LLMs was 86.2%, with individual model accuracy ranging from 79.3% to 91.7%.
- Significant differences in performance were observed among LLMs (p < 0.001).
- Microsoft Copilot (GPT-5), Google Gemini 2.5 Flash, and OpenAI o3 (Reasoning) showed superior performance compared to Perplexity AI.
Conclusions:
- LLMs demonstrate high accuracy on Board of Pharmacy Specialties certification practice questions.
- Variability in performance across LLMs and specialty domains was limited.
- Findings support further investigation of LLMs for pharmacy practice and clinical decision support, highlighting the need for domain-specific validation.
Related Concept Videos
Pharmaceutical Equivalents
273
As defined by regulatory standards, pharmaceutical equivalents require generic drug products to have identical dosage forms and chemically identical active pharmaceutical ingredients (APIs). They must adhere to compendial or applicable standards for potency, content uniformity, disintegration times, and dissolution rates. In the case of modified-release dosage forms, variations in drug content are permissible as long as the delivered amount remains consistent with the innovator drug product.
273
Pharmaceutical Alternatives: Excipients and Impurities-Related Therapeutic Nonequivalence
241
Pharmaceutical products contain more than just the active drug; they also contain various excipients such as binders, solubilizers, stabilizers, preservatives, and other elements. In some cases, impurities or contaminants might be present. Traditionally, quality control in pharmaceuticals has primarily focused on the analysis of the active drug, often overlooking the impact of these additional components. The recent issue with heparin contamination by over-sulfated chondroitin sulfate, a...
241
