Related Experiment Video
Updated: May 24, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Boosting adversarial transferability in vision-language models via multimodal feature heterogeneity
Long Chen1, Yuling Chen2,3, Zhi Ouyang1
1State Key Laboratory of Public Big Data, Guizhou University, Guiyang, 550025, China.
This study introduces a novel multimodal feature heterogeneous attack framework to improve adversarial attacks on vision-language pre-training (VLP) models in medical imaging. The framework enhances attack effectiveness and transferability, demonstrating significant improvements in robustness testing.
Area of Science:
- Artificial Intelligence
- Medical Imaging
- Computer Vision
Background:
- Vision-language pre-training (VLP) models excel in medical imaging but are susceptible to adversarial examples.
- Existing adversarial attack methods have limitations in effectiveness and transferability due to under-utilization of modal differences.
Purpose of the Study:
- To propose a novel multimodal feature heterogeneous attack (MFHA) framework to enhance adversarial attack effectiveness and transferability.
- To address the weaknesses of VLP models against adversarial examples in medical imaging.
Main Methods:
- Developed a feature heterogenization method using triplet contrastive learning (data augmentation, cross-modal/intra-modal contrastive learning).
- Implemented a cross-modal variance aggregation-based multi-domain feature perturbation method for improved transferability.
- Utilized text-guided image attacks and gradient momentum for enhanced adversarial sample generation.
Main Results:
- MFHA demonstrated a significant advantage in transferable attack capability, with an average improvement of 16.05%.
- Achieved outstanding attack performance on multimodal large language models (LLMs) like MiniGPT4 and LLaVA.
- The proposed methods effectively heterogenize consistent features into distinct ones, boosting adversarial capability.
Conclusions:
- The MFHA framework offers a robust solution for enhancing adversarial attacks on VLP models in medical imaging.
- The study highlights the importance of exploiting modal differences for effective and transferable adversarial attacks.
- The open-sourced code facilitates further research in VLP model security and robustness.
Related Concept Videos
Improving Translational Accuracy
Multi-input and Multi-variable systems
In the absence...
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Stereotype Content Model
Vision
Language and Cognition

