在对生成语言模型进行概括性临床评估时,上下文匹配不是推理
Andrew Wen1,2, Qiuhao Lu1, Yu-Neng Chuang2
1Center for Translational AI Excellence and Applications in Medicine, D. Bradley McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, TX, USA.
生成式语言模型 (GLM) 失败了临床许可证考试的基准标准,质疑它们的现实世界的准备. 为了更强大的AI评估,需要新的适应.
科学领域:
- 人工智能的人工智能
- 医疗信息学 医疗信息学
- 自然语言处理自然语言处理.
背景情况:
- 生成式语言模型 (GLM) 的临床能力通常使用从许可证考试中获得的多选题答案 (MCQA) 基准来评估.
- GLM的独特特征引发了人们对这些传统基准的有效性的担忧.
研究的目的:
- 用八个GLM验证五个MCQA基准,评估参数大小和推理能力.
- 测试MCQA概括性的关键假设:知识应用与记忆,语义一致性和无答案场景的识别.
主要方法:
- 八个GLM被用来验证五个MCQA基准.
- 快速排列被用来测试基准概括性的三个核心假设.
- 基于参数大小和推理能力来分析模型.
主要成果:
- 所有经过测试的GLM在全球范围内都无效化了支持MCQA基准标准通用性的假设.
- 较大的GLM显示出比较小的模型更具抗扰性,但对较小的模型来说,记忆仍然是一个问题.
- 所有模型都在识别没有正确答案的场景 (零答案场景) 中显著失败.
结论:
- 由于违反了基本假设,目前的MCQA基准对于评估GLM临床能力并不完全有效.
- GLMs,特别是较小的,表现出记忆能力,并与零答案场景作斗争,影响其可靠性.
- 对基准设计的调整是必要的,以便在临床环境中对GLM进行更强大,更现实的评估.
更多相关视频
10:11Portable Intermodal Preferential Looking IPL: Investigating Language Comprehension in Typically Developing Toddlers and Young Children with Autism
Published on: December 14, 2012
06:15Using the Visual World Paradigm to Study Sentence Comprehension in Mandarin-Speaking Children with Autism
Published on: October 3, 2018
相关概念视频
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Components of Language
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Improving Translational Accuracy
Improving Translational Accuracy
