MedExpQA:医学问题答案的大型语言模型的多语言基准测试
Iñigo Alonso1, Maite Oronoz1, Rodrigo Agerri1
1HiTZ Center - Ixa, University of the Basque Country UPV/EHU, Spain.
Artificial intelligence in medicine
|August 9, 2024
概括
大型语言模型在医疗人工智能方面表现有前途,但与过时的信息和幻觉作斗争. 一个新的多语言基准,MedExpQA,揭示了显著的绩效差距,特别是对于非英语语言.
科学领域:
- 人工智能在医学中的应用
- 自然语言处理自然语言处理.
- 医疗教育 技术 技术 医学教育
背景情况:
- 大型语言模型 (LLM) 展示了人工智能驱动的医疗决策支持的潜力,在医疗许可证考试中获得了高分.
- 当前的LLM面临的挑战包括过时的知识,内容幻觉,以及医学问题答案缺乏解释性.
- 现有的基准缺乏黄金标准的解释,阻碍了对LLM推理能力的评估.
研究的目的:
- 介绍MedExpQA,这是第一个用于评估医学问答LLM的多语言基准,使用医学检查数据.
- 结合医生对正确和不正确选项的黄金解释,以便进行推理评估.
- 解决非英语语言LLM基准测试被忽视的领域.
主要方法:
- 开发了MedExpQA,这是一个新的多语言基准数据集,来源于医学检查.
- 包括所有考试选项的专家撰写的黄金解释.
- 通过使用LLM进行了全面的多语言实验,包括并非Retrieval Augmented Generation (RAG) 的LLM.
主要成果:
- 在医学问答方面,LLM的成绩在英语中达到约75%的准确性.
- 除英语以外的其他语言的准确性下降了10个点,突出显示了大量的多语言差异.
- 即使使用最先进的RAG方法,整合现有的医学知识仍然具有挑战性.
结论:
- MedExpQA为评估和改进医学问答中的LLM提供了一个关键的资源.
- 目前的LLM需要取得重大进展,以满足可靠医疗应用的质量标准,特别是在多语言环境中.
- 需要进一步的研究,以有效地纳入和利用医学知识,以提高LLM绩效.
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...


