在生成AI模型中对医疗安全信息下降的纵向分析
Sonali Sharma1,2, Ahmed M Alaa3,4, Roxana Daneshjou5,6
1Department of Radiology, Faculty of Medicine, University of British Columbia, Vancouver, BC, Canada. sonali3@stanford.edu.
NPJ digital medicine
|October 2, 2025
概括
医疗生成人工智能模型,如大语言模型 (LLM) 和视觉语言模型 (VLM),显示安全性免责声明的显著减少. 这一趋势引发了人们的担忧,因为人工智能能力在医疗保健应用中取得了进展.
科学领域:
- 人工智能在医学中的应用
- 医疗信息学 医疗信息学
- 数字健康数字健康
背景情况:
- 包括LLM和VLM在内的生成AI越来越多地被用于医学图像解释和临床问题答案.
- 人工智能反应中的不准确性需要采取关键的安全措施,例如医疗免责声明.
- 人工智能在医疗保健中的不断发展能力需要强大的安全协议.
研究的目的:
- 评估大型语言模型 (LLM) 和视觉语言模型 (VLM) 输出中医疗免责声明存在的趋势.
- 评估2022年至2025年不同AI模型代的免责声明普及率是如何变化的.
- 分析对各种医疗查询和图像的响应中免责声明的使用情况.
主要方法:
- 从TIMed-Q数据集中使用500张乳房造影,500张胸部X射线,500张皮肤学图像和500个医学问题数据集生成LLM和VLM的响应.
- 量化了跨模型代 (2022-2025) 的AI产生的输出中医疗免责声明的存在和频率.
- 在指定的期间,分析了LLM和VLM的免责声明率.
主要成果:
- 在LLM输出中免责声明的存在显著下降,从2022年的26.3%降至2025年的0.97%.
- 此外,VLM免责费率也大幅下降,从2023年的19.6%降至2025年的1.05%.
- 到2025年,大多数评估的AI模型在医疗反应中缺乏免责声明.
结论:
- 人工智能医疗免责声明的下降趋势令人担忧,特别是随着人工智能模型变得越来越复杂.
- 免责声明是必要的适应性保障措施,必须与临床安全的AI能力一起发展.
- 未来的人工智能开发必须优先考虑上下文意识的安全措施,以确保在医疗保健中负责任地使用.
相关概念视频
Pharmacovigilance
1.6K
Post-marketing surveillance is a critical component of pharmaceutical regulation, often uncovering unanticipated adverse drug reactions (ADRs) once a drug is widely used over an extended period.
This process, termed pharmacovigilance, aims to detect, evaluate, and minimize harmful effects related to medication use. The data collection for pharmacovigilance depends on spontaneous reporting systems, where healthcare professionals or patients voluntarily report suspected ADRs.
In some cases, there...
This process, termed pharmacovigilance, aims to detect, evaluate, and minimize harmful effects related to medication use. The data collection for pharmacovigilance depends on spontaneous reporting systems, where healthcare professionals or patients voluntarily report suspected ADRs.
In some cases, there...
1.6K
Regression Toward the Mean
6.8K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.8K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.5K
3.5K

