摄影脊柱病评估中的通用人工智能的临床失败:诊断精度研究
Cemre Aydin1, Ozden Bedre Duygu2, Asli Beril Karakas3
1Department of Orthopedics and Traumatology, Faculty of Medicine, Ege University, 35040 Izmir, Turkey.
Medicina (Kaunas, Lithuania)
|August 28, 2025
概括
像ChatGPT和Claude 2这样的通用大型语言模型 (LLM) 在青少年特异性脊椎病 (AIS) 的摄影评估中显示出临床上不可接受的不准确性,缺乏诊断可靠性和测量准确性.
科学领域:
- 医学成像分析
- 医疗保健中的人工智能
- 整形医学
背景情况:
- 一般用途的大型语言模型 (LLM) 在没有临床验证的情况下越来越多地用于医学图像解释.
- 青少年特异性脊椎病 (AIS) 的评估依赖于精确的放射测量.
- 在AIS摄影评估中,LLM的诊断可靠性和视觉空间推理没有得到评估.
研究的目的:
- 评估ChatGPT-4o和Claude 2的诊断可靠性,以对青少年异常脊椎病 (AIS) 进行摄影评估.
- 通过临床照片确定家庭是否可以从LLM获得可靠的初步AIS评估.
- 评估AIS评估的视觉空间推理中的LLM认知忠实性.
主要方法:
- 预期诊断准确性研究 (符合STARD标准) 涉及97名青少年 (74名患有AIS,23名患有姿势不对称).
- 由两个LLM和两个骨科住院医生对放射标准进行评估的标准化临床照片 (九次视图/患者).
- 主要结果:诊断准确性 (灵敏度/特异性),科布角一致性 (林的CCC),评分器间可靠性 (科恩的 κ),布兰德-阿尔特曼的LOA.
主要成果:
- 对非AIS来说,ChatGPT的特异性为0%,Claude 2的错误阳性为78.3%.
- 系统测量误差超过了临床耐受性 (例如,胸部曲线的ChatGPT+10.74°高估,超过了800%的耐受性).
- 在LLM中,评分器之间的可靠性低于随机 (ChatGPT κ = -0.039);对胸脊曲线观察到反向一致性.
结论:
- 一般用途的LLM在摄影AIS评估中表现出临床上不可接受的不准确性,不利于临床部署.
- 灾难性的假阳性,严重的测量错误和反向诊断一致性需要紧急的监管保障措施 (例如,欧盟人工智能法).
- 无论是LLM还是摄影人类评估都不符合独立查的可靠性门; 需要特定领域的AI和3D模式.
相关概念视频
Documentation of Nursing Diagnosis
1.8K
The nurse documents nursing diagnoses and enters them into the patient record. The identified patient's nursing diagnosis is either written out with a plan of care or entered into the electronic health record.
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
1.8K
Sensitivity, Specificity, and Predicted Value
1.9K
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
1.9K
Receiver Operating Characteristic Plot
584
A ROC (Receiver Operating Characteristic) plot is a graphical tool used to assess the performance of a binary classification model by illustrating the trade-off between sensitivity (true positive rate) and specificity (false positive rate). By plotting sensitivity against 1 - specificity across various threshold settings, the ROC curve shows how well the model distinguishes between classes, with a curve closer to the top-left corner indicating a more accurate model. The area under the ROC curve...
584


