大规模验证GPT-4作为头部CT报告校对工具的可行性
Songsoo Kim1, Donghyun Kim1, Hyun Joo Shin1
1From the Departments of Biomedical Systems Informatics (S.K., Jaewoong Kim, C.H., D.Y.) and Neurology (Joonho Kim, J.Y.), Yonsei University College of Medicine, 50-1 Yonsei-ro, Seodaemun-gu, Seoul 03722, Republic of Korea; Department of Radiology, Central Draft Physical Examination Office of Military Manpower Administration, Daegu, Republic of Korea (D.K.); Department of Radiology, Research Institute of Radiological Science and Center for Clinical Imaging Data Science (H.J.S. Y.K., S.J.), and Center for Digital Health (H.J.S., D.Y.), Yongin Severance Hospital, Yonsei University College of Medicine, Yongin, Republic of Korea; Department of Radiology, Gangnam Severance Hospital, Yonsei University College of Medicine, Seoul, Republic of Korea (S.H.L.); Departments of Radiology (M.H.) and Neurology (S.J.L.), Ajou University Hospital, Ajou University School of Medicine, Suwon, Republic of Korea; and Institute for Innovation in Digital Healthcare, Severance Hospital, Seoul, Republic of Korea (D.Y.).
像GPT-4这样的大型语言模型通过检测和修改错误来提高放射学报告的准确性. 虽然GPT-4对事实错误有效,但需要进一步开发,以便在放射学中优先考虑临床意义.
科学领域:
- 医疗成像中的人工智能
- 放射学报告 分析 分析
- 在医疗保健中的自然语言处理.
背景情况:
- 放射科医生的工作负担导致了燃烧和放射学报告中的潜在错误.
- 大型语言模型 (LLM) 为医疗文档的自动错误检测和修订提供了一个潜在的解决方案.
研究的目的:
- 评估使用GPT-4用于头部CT放射学报告中的错误检测,推理和修订的可行性.
- 将GPT-4的临床实用性与人类读者进行比较,以识别和纠正报告错误.
主要方法:
- 从MIMIC-III数据集中对10,300头CT报告进行了回顾性分析.
- 实验1:评估GPT-4的错误检测,推理和修订在400个报告 (300个原始,300个带有应用错误) 上,并对200个报告进行初始优化.
- 实验2:验证了GPT-4在10,000个无错报表上的检测性能,以评估假阳性率.
主要成果:
- 在检测解释 (84%) 和事实 (89%) 错误方面,GPT-4实现了高灵敏度.
- 人类读者对事实错误的敏感性较低 (0.33-0.69),与GPT-4 (16秒) 相比,需要更长的审查时间 (82-121秒).
- 在1万份报告中,GPT-4检测到96个错误,具有较低的阳性预测值 (0.05),尽管14%的假阳性结果可能是有益的.
结论:
- GPT-4有效地检测,推理和修改放射学报告中的错误,在识别事实上的不准确性方面表现出强的表现.
- 该模型优先考虑临床重要发现的能力仍然是一个局限性.
- GPT-4提供了一个可行的工具来提高放射学报告的质量,承认其当前的优势和局限性.


