用于手部骨折检测的多模式大语言模型的诊断准确性和稳定性:在平面放射图上进行多次运行评估
Ibrahim Güler1,2, Gerrit Grieb3,4, Armin Kraus1
1Department of Plastic, Aesthetic and Hand Surgery, Otto-von-Guericke University, 39120 Magdeburg, Germany.
Diagnostics (Basel, Switzerland)
|February 13, 2026
概括
多模式大语言模型 (MLLMs) 对骨折检测有希望,但缺乏诊断稳定性. GPT-5 Pro提供了最好的准确性和一致性,尽管MLLM仍然是实验工具.
科学领域:
- 放射学 放射学是指放射学
- 人工智能的人工智能
- 医疗成像医学成像
背景情况:
- 多模式大语言模型 (MLLMs) 正在探索用于自动断裂检测.
- 在重复推断下,MLLM的诊断稳定性和一致性尚未得到充分理解.
- 本研究评估了四种领先的MLLM在检测手部骨折的性能.
研究的目的:
- 评估四种MLLM在识别手部骨折方面的诊断准确性和稳定性.
- 为了比较GPT-5 Pro,Gemini 2.5 Pro,Claude Sonnet 4.5和Mistral Medium 3.1.的模型内部的一致性和可靠性.
- 评估MLLM在不同骨折类型的表现及其对人口统计推断的能力.
主要方法:
- 四个MLLM (GPT-5 Pro,Gemini 2.5 Pro,Claude Sonnet 4.5,Mistral Medium 3.1) 分析了来自65名患者的手部放射图像.
- 每张图像被推断出五次,每个模型使用相同的零射击提示,共计1300个推断.
- 评估了诊断准确性,间运行可靠性 (Fleiss' κ) 和人口推断能力.
主要成果:
- GPT-5 Pro获得了最高的精度 (64.3%) 和一致性 (κ = 0.71),其次是双子座2.5 Pro (56.9%, κ = 0.57).
- 米斯特拉中等 3.1 显示高一致 (κ = 0.88),但精度低 (38.5%),表明"自信幻觉".
- 克劳德·索内特4.5表现出低准确度 (33.8%) 和一致性 (κ = 0.33),表明不稳定性;骨骨折对所有模型都是具有挑战性的;人口统计推断很差.
结论:
- 诊断准确性和一致性是不同的绩效指标;高的一致性并不保证正确性.
- 在评估的MLLMs中,GPT-5 Pro表现出最好的准确性和稳定性的平衡.
- 目前的MLLM是实验性诊断推理系统,而不是用于临床骨折检测的可靠独立工具.
相关概念视频
Uncertainty in Measurement: Accuracy and Precision
105.2K
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value.
105.2K
Language
921
Language is a unique communication system that uses words and systematic rules to organize and transmit information. Unlike other forms of communication, which may involve postures, movements, odors, or vocalizations, language relies on symbols and grammar. This makes human communication distinct from that of other species, who also communicate but do not use language in the same way humans do.
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
921
Improving Translational Accuracy
15.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.0K
Heart Failure IV: Classification and Diagnostic Evaluation
429
Heart failure can be classified in various ways, with the most common classifications based on physical activity limitations, disease progression, severity, and treatment strategies.The Functional Classification of Heart Failure divides patients into four categories based on physical activity limitation due to symptom burden.Class I: Patients in this class have cardiac disease but no physical activity limitations. Ordinary activities like walking, climbing stairs, or routine tasks do not cause...
429
Peripheral Arterial Disease II: Clinical Manifestations and Diagnostic Evaluation
458
Clinical manifestationsPeripheral Arterial Disease (PAD) manifests through a range of symptoms, from the characteristic intermittent claudication to atypical presentations and severe complications in advanced stages. Intermittent claudication, a hallmark symptom of PAD, presents as exercise-induced muscle pain that typically resolves within minutes of rest. This pain is reproducible and stems from inadequate blood flow, leading to the accumulation of lactic acid produced during anaerobic...
458
Irritable Bowel Syndrome II: Clinical Features and Diagnostic Evaluation
900
Irritable Bowel Syndrome II: Clinical Features and Diagnostic Evaluation
Irritable Bowel Syndrome (IBS) is classified into subtypes based on the predominant bowel habits as determined by the Bristol Stool Form Scale (BSFS). The subtypes are:
Irritable Bowel Syndrome (IBS) is classified into subtypes based on the predominant bowel habits as determined by the Bristol Stool Form Scale (BSFS). The subtypes are:
900


