对于一系列复杂的案例报告,生成性人工智能的诊断性能
Takanobu Hirosawa1, Yukinori Harada1, Kazuya Mizuta1
1Department of Diagnostic and Generalist Medicine, Dokkyo Medical University, Tochigi, Japan.
Digital health
|September 4, 2024
概括
生成型人工智能 (AI) 的诊断性能各不相同. 与Google Gemini和LLaMA2聊天机器人相比,ChatGPT-4在复杂的医疗病例中在差异诊断中表现出更高的准确性.
科学领域:
- 医疗人工智能 医疗人工智能
- 临床诊断 临床诊断 临床诊断
- 大型语言模型
背景情况:
- 使用跨不同医学专业的大型语言模型 (LLM) 的生成性人工智能 (AI) 的诊断能力在很大程度上仍未被描述.
- 评估AI诊断性能对于理解它们在临床决策支持中的潜在作用至关重要.
研究的目的:
- 评估和比较领先的生成AI的诊断性能,以生成复杂医疗病例的差异诊断.
- 为了识别不同基于LLM的AI平台之间的精度差异.
主要方法:
- 从"美国病例报告杂志" (2022年1月至2023年3月) 分析了392份已发表的病例报告,不包括儿科病例和以管理为重点的病例.
- 三个生成AI (ChatGPT-4,Google Gemini,LLaMA2聊天机器人) 从病例描述中生成了前10个差异诊断列表.
- 两个医生独立验证了最终诊断在人工智能生成的列表中是否包含.
主要成果:
- 聊天GPT-4在其十大差异诊断 (DDx) 清单中实现了最终诊断的86.7%的纳入率,明显超过了谷歌双子座 (68.6%) 和LLaMA2聊天机器人 (54.6%).
- 聊天GPT-4也显示了最高的匹配率最终诊断作为顶部列出的诊断 (54.6%),其次是谷歌双子 (31.4%) 和LLaMA2聊天机器人 (23.0%).
- 统计分析证实,ChatGPT-4的诊断准确度优于谷歌双子和LLaMA2聊天机器人 (P < 0.001),谷歌双子比LLaMA2聊天机器人优越 (前10个DDx的P < 0.001,顶级诊断的P = 0.010).
结论:
- 生成性AI在复杂的医疗病例系列中表现出不同水平的诊断性能.
- 与谷歌双子和LLaMA2聊天机器人相比,ChatGPT-4的诊断准确度更高,用于差异诊断.
- 了解这些性能差异对于将生成AI有效地整合到临床实践中,特别是一般医学中至关重要.
相关概念视频
Non-equilibrium in the Cell
4.3K
An important concept in studying metabolism and energy is that of chemical equilibrium. Most chemical reactions are reversible. They can proceed in both directions, releasing energy into their environment in one direction, and absorbing it from the environment in the other direction. The same is true for the chemical reactions involved in cell metabolism, such as the breaking down and building up of proteins into and from individual amino acids, respectively. Reactants within a closed system...
4.3K
Case Studies
11.6K
There are many research methods available to psychologists in their efforts to understand, describe, and explain behavior and the cognitive and biological processes that underlie it.
11.6K


