快速比较用于诊断并发症患者的大型语言模型:利用LLM-as-a-Judge方法进行比较研究
Peter Sarvari1, Zaid Al-Fagih1
1Rhazes AI, First Floor, 85 Great Portland Street, London, W1W 7LT, United Kingdom.
JMIRx med
|August 29, 2025
概括
在对真实患者数据进行评估的21个大型语言模型 (LLM) 中,Gemini 2.5显示出卓越的诊断准确性. 这项研究强调了LLM在改善医疗诊断方面的潜力,
科学领域:
- 医学的人工智能
- 临床决策支持系统
- 医疗保健中的自然语言处理
背景情况:
- 诊断错误导致患者死亡率很高,是美国第三大死因.
- 大型语言模型 (LLM) 有望帮助临床医生进行诊断,但缺乏对现实患者队伍的比较性能数据.
研究的目的:
- 通过使用大量真实患者数据集,比较18种流行的大型语言模型 (LLM) 的诊断能力.
- 评估不同提示和温度设置对LLM诊断性能的影响.
- 评估提取增强生成 (RAG) 在提高LLM诊断准确性的有效性.
主要方法:
- 对1000名随机选择的重症监护医疗信息中心 (MIMIC-IV) 住院患者进行了评估.
- 使用LLM-as-a-judge方法进行自动评估,将LLM生成的诊断与患者记录中的最终诊断代码进行比较.
- 使用聚合的z-测试来计算诊断命中率和评估统计意义.
主要成果:
- 在GPT-4.1的评估中,Gemini 2.5获得了最高的诊断成功率 (97.4%),超过了其他领先的模型,如GPT-4.1和Claude-4 Opus.
- 在GPT-4 Turbo的单独评估中,GPT-4.1表现最高,这表明评估的变化取决于法官LLM.
- 检索增强生成 (RAG) 显著提高了GPT-4o 05-13的命中率0.8% (P<.006),并且性能在不同的提示上有所不同.
结论:
- 在临床环境中,LLM显著提高诊断准确性的潜力.
- 用多种数据集和临床验证进行进一步的研究对于完全理解和实施LLM诊断工具至关重要.
- 人工智能开发人员和医生之间的密切合作对于LLM在医疗保健中的负责任整合至关重要.
更多相关视频
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
9.3K
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
570
相关概念视频
Classification of Illness
7.9K
The meaning of illness is individualized to each person who experiences an alteration in health. In contrast, disease is a medical term indicating a pathological change in the structure and function of the body or mind. It is a condition that has specific symptoms and boundaries.
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
7.9K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
708
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
708
Mechanistic Models: Compartment Models in Individual and Population Analysis
85
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
85
