Related Experiment Video
Updated: Jan 8, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Battle of the (Chat)Bots: Comparative Evaluation of April 2025 AI Model Accuracy on Pharmacotherapeutic Cases
Bao Anh C Tran1, Gwendolyn A Wantuch1, Moses Mangasar1
1University of South Florida, Taneja College of Pharmacy, Tampa, FL, USA.
Objective:
Generative artificial intelligence (AI) tools are increasingly utilized in health care education, offering potential learning benefits amid concerns about accuracy, reasoning, and ethical use. There is limited literature within pharmacy education, particularly quantitative studies, that explores their clinical application. This study sought to evaluate and compare the performance of 3 AI tools: ChatGPT (OpenAI), Copilot (Microsoft), and Claude AI (Anthropic), in responding to clinical patient minicases to assess their ability to apply therapeutic principles and clinical reasoning.
Methods:
Six patient minicases of varying complexities were selected from 2 integrated pharmacotherapeutic courses. Each AI model received identical prompts requiring therapeutic categorization (optimal, suboptimal, harmful) and rationale, with responses graded using existing course rubrics. Statistical analyses were conducted to compare performance across chatbots and rubric domains.
Results:
Claude AI achieved the highest average (94%, SD 9%) compared to ChatGPT (72%, SD 9.4%) and Copilot (53%, SD 18.3%), with a statistically significant difference in performance in 1 pharmacotherapeutics course. Claude AI consistently outperformed across rubric domains, whereas ChatGPT showed moderate accuracy but variable reasoning depth. Copilot had the lowest performance and highest variability. Reference quality, therapeutic plan accuracy, and defense of the therapeutic plan varied among platforms.
Conclusion:
Claude AI demonstrated the strongest overall performance. However, all tools exhibited limitations, reinforcing the need for clinical expert oversight. Though AI chatbots may support pharmacy education, our findings suggest that their current ability to perform multistep analysis remains limited. They often lack the clinical judgment, experience, and context needed for nuanced decision-making.
Related Concept Videos
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Pharmacovigilance
This process, termed pharmacovigilance, aims to detect, evaluate, and minimize harmful effects related to medication use. The data collection for pharmacovigilance depends on spontaneous reporting systems, where healthcare professionals or patients voluntarily report suspected ADRs.
In some cases, there...
Pharmacokinetic Models: Overview
There are three primary types of models: empirical, compartment, and physiological. Empirical models, with minimal...
Measurement of Bioavailability: Pharmacodynamic Methods
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...

