Related Experiment Video
Updated: Jan 17, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.0K
Performance Evaluation of 18 Generative AI Models (ChatGPT, Gemini, Claude, and Perplexity) in 2024 Japanese
Hiroyasu Sato1,2, Katsuhiko Ogasawara3,4, Hidehiko Sakurai2
1Department of Pharmacy, Abashiri-Kosei General Hospital, Abashiri, Japan.
JMIR Medical Education
|September 18, 2025
Summary
New generative artificial intelligence (AI) models show over 80% accuracy on the Japanese Pharmacist Licensing Examination, surpassing previous versions. However, AI still requires human oversight due to remaining errors in complex subjects.
Area of Science:
- Artificial Intelligence in Healthcare
- Pharmacy Education Technology
- Large Language Models
Background:
- Generative artificial intelligence (AI) is increasingly applied in healthcare.
- Previous studies assessed AI on medical exams, but pharmacy licensing exams are less explored.
- Pharmacists require diverse knowledge, necessitating AI evaluation in this field.
Purpose of the Study:
- To evaluate the performance of 18 new online chat-based large language models (OC-LLMs) on the 107th Japanese National License Examination for Pharmacists (JNLEP).
- To compare the accuracy of 2024 OC-LLMs against earlier models and identify areas for improvement.
Main Methods:
- The 107th JNLEP (345 questions) served as the benchmark.
- OC-LLMs were prompted with original Japanese questions; image uploads were used where permitted.
- Accuracy was calculated by subject area and question type, with Fleiss' κ measuring consistency.
Main Results:
- Four leading models (ChatGPT o1, Gemini 2.0 Flash, Claude 3.5 Sonnet, Perplexity Pro) achieved >80% accuracy, exceeding the passing threshold.
- Significant accuracy improvements were observed in text-only and diagram-based questions compared to earlier models.
- Accuracy for chemistry-related and chemical structure questions remained lower, with moderate consistency among top models (κ=0.334).
Conclusions:
- Online chat-based large language models (OC-LLMs) demonstrate significantly improved capabilities for pharmacy licensing examination content.
- Despite high accuracy (>80%), the remaining error rate necessitates continued human oversight in clinical practice.
- The 107th JNLEP provides a crucial benchmark for ongoing generative AI performance evaluations in pharmacy.
Related Concept Videos
Pharmacokinetic Models: Comparison and Selection Criterion
338
Physiological and compartmental models are valuable tools used in studying biological systems. These models rely on differential equations to maintain mass balance within the system, ensuring an accurate representation of the dynamic processes at play.
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
338
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.6K
3.6K
