临床医学中大型语言模型的LLM辅助系统审查
Sully F Chen1, Anton Alyakin2,3,4, Andreas Seas5
1Duke University School of Medicine, Durham, NC, USA. sully.chen@duke.edu.
Nature medicine
|March 3, 2026
概括
医学中的大型语言模型 (LLM) 正在快速增长,但严格的患者数据很少. 大多数研究使用模拟数据,强调在广泛临床采用之前需要更多的前性试验.
科学领域:
- 医疗信息学 医疗信息学
- 医疗保健中的人工智能
- 临床研究评估临床研究评估
背景情况:
- 大型语言模型 (LLM) 的临床评估自2022年以来激增.
- 研究的快速扩张对手册审查和证据综合提出了挑战.
- LLM提供了一个可扩展的解决方案,用于评估不断增长的医学文献.
研究的目的:
- 系统地审查和分析临床医学LLM应用的景观.
- 评估支持在医疗保健环境中使用LLM的证据的质量和类型.
- 确定LLM评估的趋势,包括数据源,研究设计和模型性能.
主要方法:
- 一项LLM辅助审查确定了临床医学4,609项同行评审研究 (2022年1月至2025年9月).
- 研究根据数据源 (真实世界患者数据与模拟场景),任务类型和LLM模型进行分类.
- 通过对人类表现进行头对头比较来分析表现.
主要成果:
- 只有1048项研究使用了真实世界患者数据,19项是前性随机试验.
- 模拟场景 (1,857) 和考试任务 (1,704) 主导了研究设计.
- 聊天GPT (65.7%) 和Gemini/Bard (13.1%) 是评价最多的LLM;LLM在33%的对比比较中表现优于人类.
- 25%的研究样本大小小小于30个.
结论:
- 尽管医学LLM研究的普及,强有力的,以患者为中心的证据仍然有限.
- 目前的大多数研究依赖于模拟数据或任务,而不是现实世界的临床场景.
- 较大的前性试验对于在广泛临床实施之前验证LLM的疗效和安全至关重要.
相关概念视频
Introduction to Language of Pathophysiology l
Pathophysiology investigates how biological mechanisms—typically starting at the cellular level—disrupt normal bodily functions. It bridges anatomy and physiology to explain the progression of disease. With this foundation, it is important to understand the following key terms used to describe disease processes: Diagnosis:The process of identifying a disease using clinical evaluation, including signs (objective evidence like rashes), symptoms (subjective experiences like pain), laboratory test...
Introduction to Language of Pathophysiology ll
This lesson explores key terms that describe how diseases progress, their outcomes, and their distribution in populations.Diagnostic tests identify diseases and monitor treatment. These include blood and urine tests, biopsies, imaging (X-ray, MRI), and detection of infectious agents.Remission is a reduction or disappearance of symptoms.Exacerbation refers to the worsening of symptoms, such as increased wheezing during an asthma attack.A precipitating factor triggers an acute episode, while a...


