在使用大型语言模型的随机临床试验中评估偏差风险
Honghao Lai1,2, Long Ge1,2,3, Mingyao Sun4
1Department of Health Policy and Management, School of Public Health, Lanzhou University, Lanzhou, China.
JAMA network open
|May 22, 2024
概括
大型语言模型 (LLM) 在随机临床试验 (RCT) 中评估偏差风险方面表现有前途,其准确性和一致性很高. 这些人工智能工具可以支持系统审查,尽管特定领域需要进一步改进.
科学领域:
- 医疗信息学 医疗信息学
- 临床试验方法论 临床试验方法论
- 医疗保健中的人工智能
背景情况:
- 系统性审查对于以证据为基础的医学至关重要,但需要大量的劳动力.
- 大型语言模型 (LLM) 提供了简化系统审查流程的潜力.
- 在随机临床试验 (RCT) 中评估偏差风险 (ROB) 的LLM的可靠性需要调查.
研究的目的:
- 评估在RCT中使用LLMs进行ROB评估的可行性和可靠性.
- 将两个LLM (ChatGPT和Claude) 的表现与专家评估进行比较.
主要方法:
- 一项调查研究评估了30个RCT,使用ChatGPT (LLM 1) 和Claude (LLM 2) 的结构化提示.
- 通过使用修改后的Cochrane ROB工具进行ROB评估,每个RCT都被两种模型评估两次.
- 结果与三个人类专家建立的标准标准进行了比较,计算了准确性,一致性和效率指标.
主要成果:
- 这两种LLM都表现出高的正确评估率,Claude (LLM2) 达到89.5%,ChatGPT (LLM1) 达到84.5%.
- 一致性率很高 (LLM 1的84.0%,LLM 2的87.3%),Cohen的kappa在两个模型中的大多数领域都超过了0.80.
- 克劳德 (LLM 2) 的速度明显更快,平均需要53秒,而ChatGPT (LLM 1) 的时间为77秒.
结论:
- 实际上,LLM,特别是ChatGPT和Claude,在对RCT的ROB评估中显示出相当大的准确性和一致性.
- 这些发现表明,LLM可以在系统审查过程中作为有价值的支持工具.
- 对于敏感度和F1分数较低的特定ROB领域,需要进一步关注.
相关概念视频
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
125
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
125
Hazard Ratio
114
The hazard ratio (HR) is a widely used measure in clinical trials to compare the risk of events, such as death or disease recurrence, between two groups over time. It reflects the ratio of hazard rates—the instantaneous risk of the event occurring—between a treatment group and a control group. This measure provides valuable insights into the relative effectiveness of a treatment by assessing how the risk of an event differs between the two groups.
For example, in a clinical trial...
For example, in a clinical trial...
114
Language and Cognition
342
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
342
Randomized Experiments
6.9K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
6.9K
Strategies for Assessing and Addressing Confounding
93
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
93
Improving Translational Accuracy
2.6K
2.6K


