相关实验视频
Updated: May 12, 2026

One Dimensional Turing-Like Handshake Test for Motor Intelligence
Published on: December 15, 2010
基准测试人类-人工智能合作,以获得共同的证据评估工具
Tim Woelfle1, Julian Hirt2, Perrine Janiaud3
1Pragmatic Evidence Lab, Research Center for Clinical Neuroimmunology and Neuroscience Basel (RC2NB), Basel, Switzerland; Department of Neurology, University Hospital Basel, Basel, Switzerland; Translational Imaging in Neurology (ThINk), Department of Biomedical Engineering, University Hospital and University of Basel, Basel, Switzerland.
大型语言模型 (LLM) 在评估科学证据方面显示精度低于人类. 然而,人类-人工智能协作提高了证据评估任务的准确性和效率.
科学领域:
- 科学研究中的人工智能
- 医学证据评估 医学证据评估
- 系统审查方法 系统审查方法
背景情况:
- 在证据评估中评估大型语言模型 (LLM) 的实用性对于优化研究工作流程至关重要.
- 目前评估科学报告和方法论严谨性的方法需要大量的时间和资源.
研究的目的:
- 在评估系统性审查和临床试验设计时,量化LLM与人类共识之间的协议.
- 确定人类-人工智能合作的机会,以提高证据评估的效率.
主要方法:
- 五个LLM被用来评估112个系统性审查 (使用PRISMA和AMSTAR标准) 和56个随机对照试验 (使用PRECIS-2标准).
- 在人类共识,个人人类评分器,个人LLM,联合LLM和人类-AI合作之间测量了协议.
- 当在LLM之间或在人类评级者和LLM之间出现不一致时,评级被标记为延期.
主要成果:
- 个人人类评分器在PRISMA/AMSTAR中达到89%的准确性,在PRECIS-2中达到75%的准确性.
- 单个LLM的准确率在38%至74%之间,合并LLM的准确率在64%至89%,但延期率很高.
- 人与人工智能的协作产生了最高的准确性 (80%-96%),延期率不同.
结论:
- 单独的LLM在评估科学证据方面表现比人类评估人员要差.
- 人与人工智能的协作显示出减少系统审查报告和严格评估工作负担的潜力.
- 复杂的任务,如临床试验设计评估 (PRECIS-2),并没有显著改善目前的LLM-人类合作.
更多相关视频
相关概念视频
Steady Flow of a Fluid Stream
During this process, the momentum of the fluid within the control volume remains constant over the time interval dt. By applying the...
Conservation of Mass in Moving, Nondeforming Control Volume
In the context of a detention basin, the conservation of mass states that the total mass of water entering the basin must equal the mass leaving the basin plus any accumulation of...
Typical Model Studies
Design Example: Creating a Hydraulic Model of a Dam Spillway
Gradually Varying Flow
Rapidly Varying Flow

