在回答临床研究问题时,大语言模型的准确性:系统审查和网络元分析
Ling Wang1,2, Jinglin Li2, Boyang Zhuang3
1Fuzhou University Affiliated Provincial Hospital, Shengli Clinical Medical College, Fujian Medical University, Fuzhou, China.
Journal of medical Internet research
|April 30, 2025
概括
这项研究评估了医学中的大型语言模型 (LLM),发现ChatGPT-4o在客观问题上表现出色,而ChatGPT-4在开放问题上表现出色. 人类专家在初步诊断方面仍然优越,而其他LLM在特定领域表现出优势,如前5名诊断和分拣.
科学领域:
- 医疗信息学 医疗信息学
- 医疗保健中的人工智能
- 临床决策支持系统 临床决策支持系统
背景情况:
- 大型语言模型 (LLM) 越来越多地被用于医学应用.
- 由于医学的复杂性和高风险性质,人们对LLM准确性存在担忧.
- 现有的关于医学LLM绩效的研究得出了不一致的结论.
研究的目的:
- 进行网络元分析 (NMA),评估LLM在回答临床研究问题的准确性.
- 提供高水平的证据,以指导未来的医学LLM的开发和应用.
主要方法:
- 对临床问题的答案中LLM准确性的研究进行了系统审查和NMA.
- 在2024年10月14日之前搜索了PubMed,Embase,科学网和Scopus.
- 在客观问题,开放式问题,顶部1/3/5诊断以及使用贝叶斯方法和积累排名曲线下的表面 (SUCRA) 的贝叶斯方法和分类/分类中比较了LLM准确性.
主要成果:
- 分析了168篇文章,包括35,896个问题和3063个案例; 76.2%的文章具有中等偏见风险.
- 在客观问题准确度方面,ChatGPT-4o (SUCRA=0.9207) 领先,其次是Aeyeconsult (0.9187) 和ChatGPT-4 (0.8087).
- 聊天GPT-4 (0.8708) 在开放式问题中表现出色;人类专家在诊断准确度方面排名前1 (0.9001) 和前3 (0.7126).
- 克劳德3 Opus (0.9672) 是最好的前5个诊断,而双子座 (0.9649) 在分拣和分类准确度方面领先.
结论:
- 特定的LLM表现出明显的优势:客观问题的ChatGPT-4o,开放式查询的ChatGPT-4.
- 人类临床医生在初始诊断准确度方面仍然优越 (前1名和前3名).
- 像Claude 3 Opus和Gemini这样的LLM在特定的临床任务 (前5个诊断,分类/分类) 中提供强的性能,有助于临床决策.
更多相关视频
相关概念视频
Clinical Trials
6.6K
Clinical trials are prospective experimental studies conducted on humans to determine the safety and efficacy of treatments, drugs, diet methods, and medical devices. Using statistics in clinical trials enables researchers to derive reasonable and accurate conclusions from the collected data, allowing them to make wise decisions in uncertain situations. In medical research, statistical methods are crucial for preventing errors and bias.
There are four phases in a clinical trial. A phase one...
There are four phases in a clinical trial. A phase one...
6.6K
Language and Cognition
301
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
301
Improving Translational Accuracy
2.5K
2.5K
Analysis of Population Pharmacokinetic Data
203
Analysis of population pharmacokinetic data involves studying the behavior of drugs within diverse populations to understand their pharmacokinetic parameters. Traditional pharmacokinetic methods typically involve collecting samples from a few individuals and estimating these parameters. While these methods are commonly used, they have limitations in capturing the variability in drug response among individuals or heterogeneous populations. Population pharmacokinetics is employed to address these...
203
Clinical Trials: Overview
2.6K
Clinical development focuses on how the drug will interact with the human body and encompasses four key phases of clinical trials, each serving a specific purpose in assessing the safety and effectiveness of new drugs. These phases overlap and build upon one another. Phase I involves a small group of healthy volunteers (typically 20-80 individuals) or, in cases where significant toxicity is expected, patients with the targeted disease, such as cancer or AIDS. The volunteers are tested for...
2.6K
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
106
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
106


