Related Experiment Videos
Flexible Targeted Adversarial Alignment with Frequency-based Visual Focus for Attacking Large Vision-Language Models
Abstract:
Although Large Vision-Language Models (LVLMs) have shown great capabilities across various downstream tasks, they are proven to be vulnerable to adversarial examples, thus resulting in critical safety issues. Existing adversarial attacks against LVLMs predominantly rely on uniform, image-wise perturbations, while largely overlooking how these models internally perceive and prioritize visual structures during multimodal alignment. Moreover, most prior LVLM attack methods are designed for either white-box or black-box settings, heavily depending on full-model gradients or sophisticated transfer heuristics, which incur substantial computational and resource costs and limit their practical scalability. In this paper, we propose a novel and efficient gray-box attack framework for LVLMs that only accesses the visual encoder and explicitly exploits the frequency sensitive visual focus mechanisms underlying LVLM's reasoning. Our key insight is that LVLMs exhibit non-uniform sensitivity to different frequency components of visual features, and adversarial perturbations that are selectively aligned with these frequency aware visual focuses can induce more effective and controllable targeted semantic misalignment, even without interacting with the language model or cross-modal fusion modules. Specifically, we introduce a frequency-based visual focus modeling strategy that adaptively identifies and emphasizes visually salient frequency components that are most sensitive to the LVLM's intrinsic focus, enabling the extraction of key adversarially exploitable visual features. Building upon these focused visual representations, we further propose a flexible global-subject-local adversarial alignment mechanism, which jointly aligns holistic visual semantics and localized discriminative features with the target semantic embeddings. This design allows the attack to capture both coarse-grained semantic correspondence and fine grained visual cues and thereby facilitates more robust and controllable adversarial alignment. Our framework provides a practical and effective paradigm for studying frequency-aware adversarial vulnerabilities in LVLMs under realistic gray-box settings. Extensive experiments on multiple LVLM architectures and benchmarks demonstrate that our method consistently out performs state-of-the-art attacks.