从开放式到多选项:评估ChatGPT,Google Gemini和Claude AI的诊断性能和一致性
Yaroslav O Mykhalko1, Yaroslav F Filak1, Yuliia V Dutkevych-Ivanska1
1UZHHOROD NATIONAL UNIVERSITY, UZHHOROD, UKRAINE.
概括
像Claude AI 3.5 Sonnet和ChatGPT这样的自由可用的大型语言模型 (LLM) 对差异诊断有希望,但建议对初始疾病诊断保持谨慎. 谷歌的双子座是双子座.
科学领域:
- 人工智能在医学中的应用
- 临床决策支持系统 临床决策支持系统
- 在医疗保健中的自然语言处理.
背景情况:
- 大型语言模型 (LLM) 越来越多地被用于医学应用.
- 评估LLMs的诊断准确性和可靠性对于临床整合至关重要.
- 自由可用的LLM为医疗保健专业人员提供了潜在的可访问性.
研究的目的:
- 评估著名的LLMs的诊断性能和响应可重复性.
- 为了比较ChatGPT 3.5,ChatGPT 4o,谷歌双子座和克劳德AI 3.5 Sonnet.
- 用详细的临床病例描述来评估LLM的疗效.
主要方法:
- 用100个临床病例描述来测试四个LLM.
- 评估有两个阶段:第一阶段 (仅描述),第二阶段 (描述和答案变体).
- 用协议百分比和科恩的卡帕 (k) 来衡量响应的一致性.
主要成果:
- 克劳德AI 3.5 索内特表现出最高的疗效 (72.00% 在第一阶段, 89.00% 在第二阶段).
- 聊天GPT模型 (3.5和4o) 显示出强的性能,特别是在第二阶段 (90.00%和95.00%的疗效).
- 谷歌Gemini的疗效较低 (第一阶段为44.00%,第二阶段为65.00%).
- 所有LLM都显示出高响应的一致性 (一致性>93%,k>0.85).
结论:
- 克劳德AI 3.5 索内特和聊天GPT模型是差异诊断的有效工具.
- 在使用LLM来从头开始诊断疾病时,建议谨慎使用.
- 谷歌双子目前的低效率引发了关于其临床可行性的疑问.
相关概念视频
Improving Translational Accuracy
9.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.0K
Detection of Gross Error: The Q Test
5.6K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
5.6K


