新开发的大型语言模型在重症病例中的诊断性能:一项比较研究
Xintong Wu1, Yu Huang1, Qing He1
1Department of Intensive Care Medicine, Affiliated Hospital of Southwest Jiaotong University, The Third People's Hospital of Chengdu, Chengdu, Sichuan, China.
International journal of medical informatics
|August 27, 2025
概括
新开发的大型语言模型 (LLM) 对重症监护室 (ICU) 的临床决策支持有希望. 在诊断准确度方面,ChatGPT-o3领先,而开源的DeepSeek-R1也具有竞争力.
科学领域:
- 医学的人工智能
- 临床决策支持系统
- 危急护理医学
背景情况:
- 大型语言模型 (LLM) 显示出临床决策支持的潜力.
- 在重症监护室 (ICU) 的诊断性能还未得到充分研究.
- 这项研究评估了新开发的临床临床诊断方法.
研究的目的:
- 评估四种新型LLM在重症病例中的诊断准确性.
- 将这些LLM的差异诊断质量和响应质量进行比较.
- 确定最有效的临床决策支持的临床决策工具.
主要方法:
- 一项横截面比较研究评估了四种LLM:ChatGPT-4o,ChatGPT-o3,DeepSeek-V3和DeepSeek-R1.
- 在已发表的ICU文献中对50例严重病例进行了LLM测试.
- 诊断准确性,差异性诊断质量和反应质量进行了比较.
主要成果:
- ChatGPT-o3获得了最高的诊断准确率 (72%),其次是DeepSeek-R1 (68%) 和ChatGPT-4o (64%). 据报道,DeepSeek-V3的准确率达到了32%.
- ChatGPT-o3,DeepSeek-R1和ChatGPT-4o的表现明显超过了DeepSeek-V3的表现.
- 所有评估的LLM都显示出高响应质量 (完整性,清晰性,有用性),ChatGPT- o3和DeepSeek- R1显示出优异的差异诊断质量.
结论:
- 新开发的LLM,特别是推理模型,显示出在重症监护中支持诊断的巨大潜力.
- 特定领域的微调可以进一步提高LLM诊断的准确性.
- 开源的DeepSeek-R1的竞争性表现突出了其在资源有限的环境中的潜力.
相关概念视频
Classification of Illness
7.9K
The meaning of illness is individualized to each person who experiences an alteration in health. In contrast, disease is a medical term indicating a pathological change in the structure and function of the body or mind. It is a condition that has specific symptoms and boundaries.
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
7.9K
Language and Cognition
438
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
438
Improving Translational Accuracy
11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K


