通过GPT-3.5,GPT-4和谷歌Bard:多语言研究的BI-RADS类别分配
Andrea Cozzi1, Katja Pinker1, Andri Hidber1
1From the Imaging Institute of Southern Switzerland (IIMSI), Ente Ospedaliero Cantonale, Via Tesserete 46, 6900 Lugano, Switzerland (A.C., L.B., M.C., S.R., F.D.G., S.S.); Breast Imaging Service, Department of Radiology, Memorial Sloan Kettering Cancer Center, New York, NY (K.P., R.L.G., B.C.); Faculty of Biomedical Sciences, Università della Svizzera Italiana, Lugano, Switzerland (A.H., S.R., F.D.G., S.S.); Department of Radiology, Netherlands Cancer Institute, Amsterdam, the Netherlands (T.Z., R.M.M.); Department of Diagnostic Imaging, Radboud University Medical Center, Nijmegen, the Netherlands (T.Z., R.M.M.); and GROW Research Institute for Oncology and Reproduction, Maastricht University, Maastricht, the Netherlands (T.Z.).
大型语言模型 (LLM) 在乳腺成像报告和数据系统 (BI-RADS) 类别中显示了与人类读者的中度一致. 然而,LLM产生了显著不一致的BI-RADS类别,可能会对临床管理产生负面影响.
科学领域:
- 放射学 放射学是一门学科.
- 人工智能的人工智能
- 医疗信息学 医疗信息学
背景情况:
- 大型语言模型 (LLM) 对复杂的医疗任务 (如乳腺成像解释) 的临床实用性尚未得到充分证实.
- 评估LLM在分配标准化类别方面的表现对于安全的临床整合至关重要.
研究的目的:
- 评估人类放射科医生和LLM (GPT-3.5,GPT-4,Gemini) 在赋予乳房成像报告和数据系统 (BI-RADS) 类别方面的协议.
- 确定人类和LLM分配的BI-RADS类别之间的差异对临床管理的影响.
主要方法:
- 追溯分析2400个多语言乳房成像报告 (意大利语,英语,荷兰语) 的BI-RADS类别1-5.
- 由董事会认证的放射科医生与LLM使用Gwet的协议系数 (AC1) 的BI-RADS分配的比较.
- 分析影响临床管理的类别变化及其潜在的负面影响.
主要成果:
- 人类读者达到了几乎完美的一致 (AC1=0.91).
- 在LLM中,与人类读者之间存在中度的一致性 (GPT-4:AC1=0.52,GPT-3.5:AC1=0.48,双子座:AC1=0.42).
- 与人类读者 (4,9%) 相比,LLM产生了影响临床管理的显著更多不一致的BI-RADS类别 (18.1%-25.5%),负面影响更高 (10.6%-18.1%对1.5%).
结论:
- 在不同语言中,LLM对人类的BI-RADS分类表现出适度的同意.
- LLM产生了大量不一致的BI-RADS类别,可能会对患者的护理产生不利影响.
- 在临床部署LLMs用于乳腺成像报告分析之前,需要进一步的研究和验证.
更多相关视频
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
09:21Human Brown Adipose Tissue Depots Automatically Segmented by Positron Emission Tomography/Computed Tomography and Registered Magnetic Resonance Images
Published on: February 18, 2015
