在2024年对18个生成性AI模型 (ChatGPT,Gemini,Claude和Perplexity) 的性能评估日本药剂师执照考试:比较研究
Hiroyasu Sato1,2, Katsuhiko Ogasawara3,4, Hidehiko Sakurai2
1Department of Pharmacy, Abashiri-Kosei General Hospital, Abashiri, Japan.
JMIR medical education
|September 18, 2025
概括
新的生成人工智能 (AI) 模型在日本药剂师执照考试中显示出超过80%的准确性,超过了以前的版本. 然而,由于复杂主题中的剩余错误,人工智能仍然需要人类监督.
科学领域:
- 医疗保健中的人工智能
- 药房 教育 技术 技术
- 大型语言模型
背景情况:
- 生成型人工智能 (AI) 越来越多地应用于医疗保健领域.
- 之前的研究评估了医疗考试中的AI,但药房许可证考试的探索较少.
- 药剂师需要多样化的知识,需要在这个领域进行人工智能评估.
研究的目的:
- 评估18个基于在线聊天的新型大型语言模型 (OC-LLMs) 在第107届日本药剂师执照全国考试 (JNLEP) 的表现.
- 将2024年OC-LLM与早期模型的准确性进行比较,并确定需要改进的领域.
主要方法:
- 第107届JNLEP (345个问题) 作为基准.
- OC-LLM被提示以原始的日语问题;图像上传被允许使用.
- 精度是根据学科领域和问题类型计算的,Fleiss' κ测量一致性.
主要成果:
- 四个领先的模型 (ChatGPT o1,Gemini 2.0 Flash,Claude 3.5 Sonnet,Perplexity Pro) 实现了>80%的准确性,超过了通过值.
- 与先前的模型相比,仅基于文本和基于图表的问题的准确性得到了显著改善.
- 化学相关和化学结构问题的准确性仍然较低,在顶级模型中保持适度一致 (κ=0.334).
结论:
- 基于在线聊天的大型语言模型 (OC-LLMs) 显示了药房许可证考试内容的显著改进能力.
- 尽管精度很高 (>80%),但剩余的错误率需要在临床实践中继续对人类进行监督.
- 第107届JNLEP为正在进行的制造性AI药房绩效评估提供了关键的基准.
相关概念视频
Pharmacokinetic Models: Comparison and Selection Criterion
338
Physiological and compartmental models are valuable tools used in studying biological systems. These models rely on differential equations to maintain mass balance within the system, ensuring an accurate representation of the dynamic processes at play.
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
338
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.6K
3.6K

