在急诊室GPT辅助差异诊断的准确性评估
Fatemeh Shah-Mohammadi1, Joseph Finkelstein1
1Department of Biomedical Informatics, School of Medicine, University of Utah, Salt Lake City, UT 84112, USA.
Diagnostics (Basel, Switzerland)
|August 29, 2024
概括
像ChatGPT-3.5和ChatGPT-4这样的大型语言模型可以通过分析患者笔记来协助紧急诊断. GPT-4显示出更高的准确性,特别是在复杂的病例中,帮助诊断过程.
科学领域:
- 医疗信息学 医疗信息学
- 医疗保健中的人工智能
- 紧急医疗 紧急医疗
背景情况:
- 在急诊室 (ED) 中,准确和及时的诊断对于患者的治疗结果和医疗保健效率至关重要.
- 电子健康记录 (EHR) 包含有价值的患者数据,但通常以非结构化的格式.
- 大型语言模型 (LLM) 提供了处理非结构化临床文本的潜力.
研究的目的:
- 为了评估ChatGPT-3.5和ChatGPT-4的诊断准确性,在产生ED电子病历笔记的差异诊断时.
- 为了比较不同 ChatGPT 代的性能,作为 ED 环境中的潜在诊断辅助工具.
- 将LLM生成的诊断与实际的出院诊断进行比较.
主要方法:
- 使用了从急诊室入院前24小时的电子健康记录笔记.
- 使用ChatGPT-3.5和ChatGPT-4生成差异诊断列表.
- 通过将LLM生成的列表与身体系统和类别级别的确诊出院诊断进行比较来评估准确性.
主要成果:
- 聊天GPT-3.5和聊天GPT-4都在预测身体系统水平的诊断方面表现出合理的准确性.
- 在身体系统层面上,ChatGPT-4的性能略高于ChatGPT-3.5.
- 在更精细的类别层面,性能精度下降,尽管ChatGPT-4在关键类别中显示出更好的精度.
结论:
- 在紧急医疗中,LLM是有前途的诊断辅助工具,特别是在处理非结构化EHR数据时.
- 与ChatGPT-3.5相比,ChatGPT-4显示了增强的功能,特别是在复杂的临床场景中.
- 需要进一步改进,以提高细粒度水平的诊断精度,以便在临床上广泛采用.
相关概念视频
Sensitivity, Specificity, and Predicted Value
225
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
225
Pneumonia III: Complications and Assessment
189
Pneumonia poses the potential for numerous complications that warrant consideration. These complications include the following:
189


