Related Experiment Video
Updated: Jan 8, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Diagnostic performance of four AI tools in pharmacology MCQs: Accuracy, sensitivity, and specificity
Ayah J Al-Rahahleh1, Mai Z Rizik2, Fahmi Y Al-Ashwal3,4
1Clinical Pharmacy and Therapeutics Department, Faculty of Pharmacy, Applied Science Private University, Amman, Jordan.
Background:
The rapid rise of AI in medical and pharmaceutical education has engendered much interest; however, a knowledge gap still exists in the evaluation of performances of these tools in critical academic contexts.
Objectives:
The aim of this study was to assess and compare the performances of four openly accessible AI language tools, Microsoft Copilot, ChatGPT-3.5, Google Gemini, and DeepSeek AI, in responding to pharmacology-related MCQs with regard to diagnostic accuracy, sensitivity, specificity, and reproducibility.
Methods:
A total of 80 MCQs were generated and validated, representing four therapeutic systems: cardiovascular, respiratory, gastrointestinal, and endocrine, including four pharmacological domains: mechanism of action, side effects, pharmacokinetics, and drug-drug interactions. Answers were classified into true/false positives and negatives in order to calculate accuracy, sensitivity, and specificity. After two weeks, a second round of testing was performed with the questions to assess answer reproducibility.
Results:
The top overall performer was Microsoft Copilot: 87.5% accuracy, a sensitivity of 94.6%, and a specificity of 70.8%. It continued to perform strongly across all therapeutic systems, especially in the cardiovascular and respiratory domains, with the highest accuracy in identifying drug mechanisms and side effects. ChatGPT-3.5 performed similarly to Google Gemini (76.3% and 75.0% accuracy, respectively) but with higher sensitivity for ChatGPT-3.5 and higher specificity for Gemini. DeepSeek AI had the lowest accuracy overall (68.8%) and the lowest specificity (29.2%), but the highest consistency of reproducibility (97.5%). The performance of all tools decreased significantly with increasing level of question difficulty (p < 0.05).
Conclusion:
All tools have some value in pharmacology education, but Microsoft Copilot was the most consistently accurate. Limitations in complexity and reproducibility suggest that caution should be exercised in academic and clinical use, particularly given the variability seen with ChatGPT-3.5.
More Related Videos
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
04:57Comparative Analysis of Automatic Fecal Analyzer versus Direct Wet Smear Microscopy for Detecting Parasitic Infections in Stool Samples
Published on: April 25, 2025
Related Concept Videos
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Receiver Operating Characteristic Plot
Measurement of Bioavailability: Pharmacodynamic Methods
Drug Product Performance: In Vitro–In Vivo Correlation
Therapeutic Drug Monitoring: Drug Analysis Methods
Effect of Hepatic Disease on Pharmacokinetics: Pathophysiologic Assessment and Liver Function Test