敌对的即时和微调攻击威胁到医学的大型语言模型
Yifan Yang1,2, Qiao Jin1, Furong Huang2
1National Library of Medicine (NLM), National Institutes of Health (NIH), Bethesda, MD, USA.
Nature communications
|October 9, 2025
概括
医疗保健中的大型语言模型 (LLM) 容易受到诸如即时注射和中毒数据之类的对抗性攻击. 在攻击后检测模型重量的变化是开发安全医疗AI防御的关键.
科学领域:
- 人工智能在医学中的应用
- 医疗保健中的网络安全
- 自然语言处理自然语言处理.
背景情况:
- 大型语言模型 (LLM) 显示出在改变医疗保健,帮助诊断和患者护理方面的潜力.
- 然而,LLM容易受到对抗性攻击,在敏感的医疗应用中构成风险.
研究的目的:
- 调查LLM对提示注入和数据中毒攻击的脆弱性.
- 评估医疗任务的LLM安全性,包括疾病预防,诊断和治疗.
主要方法:
- 利用现实世界的患者数据来测试开源和专有LLMs.
- 模拟敌对攻击,包括即时注射和微调有毒样本.
主要成果:
- 在各种医疗任务中证明了LLM对恶意操纵的显著脆弱性.
- 由于中毒数据,观察到微妙的微调模型重量的变化,即使没有重大性能降低.
结论:
- 医疗保健中的LLM需要强大的安全措施来防止对抗性攻击.
- 识别重量转移提供了一种潜在的方法来检测和减轻对医疗AI系统的攻击.
相关概念视频
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.5K
3.5K


