在放射学报告分类的大型语言模型中评估提示和数据干扰灵敏度
Vera Sorin1, Jeremy D Collins1, Alex K Bratt1
1Department of Radiology, Mayo Clinic College of Medicine and Science, Mayo Clinic, Rochester, MN 55905, United States.
JAMIA open
|August 13, 2025
概括
大型语言模型 (LLM) 在放射学报告中对肺栓塞 (PE) 的分类具有很高的准确性. 快速设计和数据质量显著影响LLM的性能,需要对临床使用进行仔细的验证.
科学领域:
- 医疗成像中的人工智能
- 在医疗保健中的自然语言处理.
- 放射学报告 分析 分析
背景情况:
- 大型语言模型 (LLM) 越来越多地被用于医疗保健应用,特别是在自然语言处理任务中.
- 医疗报告的准确分类,例如肺栓塞 (PE) 的医疗报告,对于患者的护理至关重要.
- 了解LLM在高风险的临床场景中的性能限制是必不可少的.
研究的目的:
- 评估谷歌大型语言模型 (LLM) 在分类肺CT血管造影放射学报告中肺栓塞 (PE) 存在的性能.
- 评估各种条件,包括不同的提示设计和数据干扰,对LLM分类准确性的影响.
- 确定在不同配置和云环境中LLM性能的稳定性和可变性.
主要方法:
- 追溯分析了11999份肺CT血管造影放射学报告.
- 三个谷歌LLM的评估:双子-1.5-Pro,双子-1.5-闪光-001和双子-1.5-闪光-002.
- 通过基于计算机视觉的PE检测 (CVPED) 算法和多次LLM运行之间的一致性来确定基本真相,并对差异进行手动审查.
- 分析快速设计,数据扰动和地理云区域对性能指标的影响.
主要成果:
- 整个LLM的整体准确性在0.953到0.996.99之间.
- 一个修改后的提示实现了回忆率高达0.997.
- 几次射击提示提高了回忆力 (高达0.99),而链式思维提示通常会降低性能.
- 双子座-1.5-Flash-002显示了对数据扰动的最高稳定性;双子座-1.5+-Pro具有最小的地理变异性,而Flash模型是稳定的.
结论:
- 在对肺栓塞的放射学报告进行分类时,LLM 显示出高效率.
- 性能对提示工程和数据质量敏感,突出了严格验证的必要性.
- 在在临床决策过程中部署LLM之前,系统评估至关重要.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
578
07:15Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
6.9K
相关概念视频
Sensitivity, Specificity, and Predicted Value
661
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
661
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Survival Tree
159
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
159
