在系统性审查中整合大型语言模型:使用ROBINS-I进行偏差评估风险的框架和案例研究
Bashar Hasan1,2, Samer Saadi3,2, Noora S Rajjoub3
1Kern Center for the Science of Healthcare Delivery, Mayo Clinic, Rochester, Minnesota, USA Hasan.Bashar@mayo.edu.
BMJ evidence-based medicine
|February 21, 2024
概括
大型语言模型 (LLM) 显示,在系统性审查中对偏见风险评估方面,与人类审稿人之间存在温和的共识. 拟议的框架指导了LLM的整合,但人类监督仍然是必不可少的.
科学领域:
- 医疗信息学 医疗信息学
- 医疗保健中的人工智能
- 证据综合 证据综合
背景情况:
- 系统性审查对于以证据为基础的医学至关重要,但耗时.
- 大型语言模型 (LLM) 具有加速系统审查流程的潜力.
- 缺乏将LLM整合到系统审查中的明确方法.
研究的目的:
- 用ROBINS-I工具评估GPT-4与人类审查员在偏差评估风险方面的协议.
- 提出一个框架,将LLMs纳入系统审查.
- 在系统审查中确定适合LLM应用的关键任务.
主要方法:
- 一个案例研究,将GPT-4的偏差评估风险与使用ROBINS-I工具的人类审查员进行比较.
- 在ROBINS-I领域使用原始百分比和肯德尔系数分析协议.
- 开发一个四个领域的框架,以在系统审查中整合LLM.
主要成果:
- 在"干预分类"领域观察到GPT-4和人类审稿人之间最高的原始协议.
- 在"参与者选择"",缺失数据"和"结果测量"领域发现了适度的协议 (肯达尔系数).
- 关于偏差风险的总体共识为61% (肯德尔系数=0.35),表明需要人类-AI合作.
结论:
- 法律法规证明了在系统审查中提供帮助的潜力,特别是在偏见风险评估方面.
- 提出了一个结构化的框架,以指导LLMs在系统审查工作流程中的整合.
- 人类监督仍然至关重要,需要将人工智能与独立的人类审查员配对,以获得可靠的结果.
相关概念视频
Strategies for Assessing and Addressing Confounding
100
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
100
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
126
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
126
Stereotype Content Model
14.7K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.7K


