大型语言模型在认知能力测试中的表现:在德国医学院入学考试 (TMS) 上进行多次运行评估
Henrik Stelling1, Armin Kraus2, Gerrit Grieb3,4
1Practices for Nuclear Medicine, Rubensstraße 125, 12157 Berlin, Germany.
European journal of investigation in health, psychology and education
|February 26, 2026
概括
大型语言模型 (LLM) 在医学能力测试中显示出有限的推理能力,尽管在知识考试中表现强. 他们的表现在不同的认知任务和重复评估中存在显著差异.
科学领域:
- 医疗教育中的人工智能
- 认知能力评估 认知能力评估
- 大型语言模型 (LLM)
背景情况:
- 法学士在基于知识的医学考试中表现出色,但他们的推理和抽象技能被理解得更少.
- 医学研究考试 (TMS) 评估认知能力,对医学院入学至关重要.
- 在TMS上评估LLM可以了解他们对医学推理的能力.
研究的目的:
- 评估TMS上各种LLM的性能和一致性.
- 为了比较基于文本的与视觉分析的认知任务的LLM绩效.
- 在多次评估中调查LLM的可靠性.
主要方法:
- 在标准化的TMS项目上测试了8个专有和开源的LLM.
- 评估涵盖了基于文本和视觉分析的认知领域.
- 采用多运行设计来评估运行间可靠性.
主要成果:
- 在TMS上LLM的准确性大大低于基于知识的考试.
- 在基于文本和视觉分析的子测试之间观察到显著的绩效差异.
- 开源LLM的性能与专有模型相比,具有异质的运行间可靠性.
结论:
- 目前的LLM在医学院招生任务方面表现出有限的,特定领域的适应性.
- 在知识回忆方面的高性能并不能保证稳定的流体智能.
- 法学士评估需要差异化,多运行策略,由于模式依赖的性能和可变性.
相关概念视频
Multiple Comparison Tests
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Measures of Intelligence
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...


