评估人工智能大型语言模型在跨语言检测儿科药物错误方面的性能:一项比较研究
Rana K Abu-Farha1, Haneen Abuzaid2, Jena Alalawneh3
1Clinical Pharmacy and Therapeutics Department, Faculty of Pharmacy, Applied Science Private University, Amman 11937, Jordan.
Journal of clinical medicine
|January 10, 2026
概括
在四个AI模型中,微软Copilot在检测儿科药物错误方面显示出最高的准确性. 性能因语言而异,阿拉伯语的准确性通常较低,强调需要更好的多语言AI培训.
科学领域:
- 医疗保健中的人工智能
- 药物监督 药物监督 药物监督
- 儿科药物安全 儿科药物安全
背景情况:
- 药物错误在儿科药物治疗中构成重大风险.
- 评估人工智能 (AI) 工具对药物错误检测的有效性至关重要.
研究的目的:
- 评估四个人工智能模型 (GPT-5,GPT-4,微软Copilot,谷歌Gemini) 在儿童病例情景中识别药物错误的性能.
- 在英语和阿拉伯语言中比较AI模型性能.
主要方法:
- 分析了60个儿科病例,其中一半包含四种治疗系统中的药物错误.
- 人工智能模型使用英语和阿拉伯语的统一提示符进行了测试.
- 性能指标包括精度,灵敏度,特异性和可重复性,使用SPSS版本22进行分析.
主要成果:
- 微软Copilot取得了最高的准确度 (86.7%的英语,85.0%的阿拉伯语),其次是GPT-5.
- 谷歌双子表现出最低的准确性 (76.7%的英语,73.3%的阿拉伯语).
- 阿拉伯语言的表现通常低于英语;微软Copilot显示出卓越的可复制性和语言间的协议.
结论:
- 在这项研究中,微软Copilot在检测儿科药物错误方面表现优于其他AI模型.
- 这些发现强调需要加强多语言AI培训,以确保跨语言的公平表现.
- 人类监督和特定领域的人工智能培训对于在儿科药物治疗中安全应用至关重要.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
1.3K
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
7.9K
相关概念视频
Improving Translational Accuracy
3.5K
3.5K
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
