在视觉语言模型中通过多式联络特征异质性提升对抗性可转移性
Long Chen1, Yuling Chen2,3, Zhi Ouyang1
1State Key Laboratory of Public Big Data, Guizhou University, Guiyang, 550025, China.
Scientific reports
|March 2, 2025
概括
这项研究引入了一种新的多式联络功能异质攻击框架,以改善对视觉语言预训练 (VLP) 模型在医学成像中的对抗性攻击. 该框架增强了攻击的有效性和可转移性,在稳定性测试中显示出显著的改进.
科学领域:
- 人工智能的人工智能
- 医疗成像医学成像
- 计算机视觉 计算机视觉
背景情况:
- 视觉语言预训练 (VLP) 模型在医学成像方面表现出色,但易受对抗性示例的影响.
- 现有的对抗性攻击方法在有效性和可转移性方面存在局限性,原因是不充分利用模式差异.
研究的目的:
- 提出一种新的多式联运特征异质攻击 (MFHA) 框架,以提高对抗性攻击的有效性和可转移性.
- 针对医学成像中的对抗性例子,解决VLP模型的弱点.
主要方法:
- 开发了一种使用三重对比学习 (数据增强,跨模式/内部模式对比学习) 的特征异质化方法.
- 实施了基于跨模态变异聚合的多域特征扰动方法,以改善可转移性.
- 利用文本导向图像攻击和梯度势头来增强对抗样本生成.
主要成果:
- 在可转移的攻击能力方面,MFHA表现出了显著的优势,平均提高了16.05%.
- 在MiniGPT4和LLaVA等多式大型语言模型 (LLM) 上实现了卓越的攻击性能.
- 提出的方法有效地将一致的特征变异为不同的特征,提高对抗能力.
结论:
- 在医疗成像中,MFHA框架提供了一个强大的解决方案,以加强对医学成像中的VLP模型的对抗性攻击.
- 该研究强调了利用模式差异对有效和可转移的对抗性攻击的重要性.
- 开源代码有助于进一步研究VLP模型的安全性和稳定性.
相关概念视频
Improving Translational Accuracy
2.5K
2.5K
Multi-input and Multi-variable systems
93
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
93
Generalization, Discrimination, and Extinction
399
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
399
Stereotype Content Model
13.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
13.9K
Vision
52.9K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
52.9K
Language and Cognition
313
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
313


