Related Experiment Video
Updated: Jul 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
MDAPT: Multi-Modal Depth Adversarial Prompt Tuning to Enhance the Adversarial Robustness of Visual Language Models.
Chao Li1, Yonghao Liao1, Caichang Ding2
1School of Computer Science, Hubei University of Technology, Wuhan 430068, China.
This study introduces Multi-modal Depth Adversarial Prompt Tuning (MDAPT) to enhance visual language models (VLMs) against adversarial attacks. MDAPT significantly boosts VLM accuracy and robustness, outperforming traditional methods.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Natural Language Processing
Background:
- Large visual language models (VLMs) like CLIP excel in performance but are vulnerable to adversarial examples.
- Existing methods often struggle to maintain robustness against sophisticated attacks.
Purpose of the Study:
- To investigate the accuracy and robustness of VLMs from a multi-modal perspective.
- To propose a novel fine-tuning method to enhance VLM resilience against adversarial attacks.
Main Methods:
- Developed Multi-modal Depth Adversarial Prompt Tuning (MDAPT), a fine-tuning technique.
- MDAPT guides visual prompt generation using text prompts for improved VLM performance.
- Conducted extensive experiments on three datasets under adversarial conditions.
Main Results:
- Achieved significant performance improvements on three datasets (ϵ=4/255).
- MDAPT increased accuracy by an average of 17.84% and robustness by 10.85% compared to manual prompts.
- Demonstrated substantial gains (32.16% accuracy, 21.00% robustness) under three different attack methods with efficient settings.
Conclusions:
- MDAPT effectively enhances the accuracy and robustness of VLMs against adversarial examples.
- The proposed multi-modal approach offers a promising direction for developing more resilient AI systems.
- MDAPT provides a significant advantage over traditional prompt design strategies in adversarial settings.
Related Concept Videos
Improving Translational Accuracy
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
Multi-input and Multi-variable systems
In the absence...
Depth Perception and Spatial Vision
Associative Learning
Classical conditioning, also known...
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...