大型语言模型对诊断推理的影响:一项随机临床试验
Ethan Goh1,2, Robert Gallo3, Jason Hom4
1Stanford Center for Biomedical Informatics Research, Stanford University, Stanford, California.
JAMA network open
|October 28, 2024
概括
大型语言模型 (LLM) 在临床试验中没有显著改善医生的诊断推理. 然而,单独LLM的表现优于医生,这表明未来AI-医生合作的潜力.
科学领域:
- 医疗人工智能 医疗人工智能
- 临床决策支持系统 临床决策支持系统
- 医生绩效评价 医生的绩效评价
背景情况:
- 大型语言模型 (LLM) 在医学推理评估中表现有前途.
- 在实际的医生诊断推理LLMs的影响仍然不清楚.
研究的目的:
- 评估大型语言模型 (LLM) 对医生诊断推理的影响,与传统资源相比.
主要方法:
- 一个单盲随机临床试验涉及多个机构的50名医生.
- 参与者被随机分配到使用传统资源或仅使用传统资源的LLM.
- 用标准化标题和专家共识来评估诊断性能.
主要成果:
- 在LLM组和传统资源组之间,诊断推理得分没有显著差异 (76%对74%).
- 每个病例所花费的时间在各组之间没有显著差异.
- 当单独使用LLM时,比传统资源组得分明显高 (16个百分点).
结论:
- 在这项研究中,将LLM作为诊断辅助工具整合并没有增强医生的临床推理.
- 仅LLM就表现出了卓越的表现,突出了需要进一步发展人工智能与医生的合作.
- 未来的研究应该专注于优化人工智能和临床实践之间的协同作用.
更多相关视频
相关概念视频
Language and Cognition
329
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
329
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
121
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
121
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Hazard Ratio
91
The hazard ratio (HR) is a widely used measure in clinical trials to compare the risk of events, such as death or disease recurrence, between two groups over time. It reflects the ratio of hazard rates—the instantaneous risk of the event occurring—between a treatment group and a control group. This measure provides valuable insights into the relative effectiveness of a treatment by assessing how the risk of an event differs between the two groups.
For example, in a clinical trial...
For example, in a clinical trial...
91
Blinding
2.4K
Blinding is a commonly used method of not telling participants which treatment a subject is receiving. Blinding is a critical part of a randomized control trial or RCT. It reduces the bias that affects the results. In an RCT, blinding is used in the form of a placebo. A placebo effect occurs when untreated subjects falsely believe they have received the treatment and report improved symptoms. A placebo or a dummy treatment is administered to subjects to negate the bias caused by such an effect.
2.4K
What is an Experiment?
10.4K
An experiment is a planned activity carried out under controlled conditions. The purpose of an experiment is to investigate the relationship between two variables. When one variable causes change in another, we call the first variable the explanatory or independent variable. The affected variable is called the response or dependent variable. In a randomized experiment, the researcher manipulates values of the explanatory variable and measures the resulting changes in the response variable. The...
10.4K


