Related Experiment Video
Updated: May 5, 2026

Advanced Diffusion Imaging in The Hippocampus of Rats with Mild Traumatic Brain Injury
Published on: August 14, 2019
Adversarial discriminant attack on text-to-image diffusion models
Hanxiao Wu1, Shengwu Xiong2, Dong Yi3
1School of Computer Science and Artificial Intelligence, Wuhan University of Technology, Wuhan, 430070, Hubei, China; Institute of Automation, Chinese Academy of Sciences, Beijing, 100190, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, 101408, China; Wuhan AI Research, Wuhan, 430000, Hubei, China.
Abstract:
Despite advancements in concept-erased diffusion models, the persistent risk of generating Not-Safe-For-Work (NSFW) content in text-to-image tasks remains a critical challenge. To expose vulnerabilities in these models, some existing works designs attack method from generation perspective, which attempts to constraint the similarity between generate images and specific inappropriate images. However, generating visually similar images does not necessarily imply that the NSFW content has been successfully reconstructed, so the effectiveness of existing attack methods remains limited. To address this limitation, we propose Adversarial Discriminant Attack (ADAtk), a novel method designed to expose vulnerabilities in concept-erased diffusion models. Unlike existing attacks that focus on generation, ADAtk adopts a more intuitive discriminative perspective, aiming to generate images that are classified as inappropriate. By optimizing the likelihood of producing NSFW content, ADAtk crafts adversarial perturbations in the model's latent space, thereby guiding the reconstruction of NSFW concepts (e.g., nudity) aligned with the target discriminant class. Experimental results show that ADAtk can achieve an over 90% success rate in bypassing current internal security mechanisms, exposing critical limitations in existing concept-erasure techniques. These findings provide essential insights for improving the safety and reliability of text-to-image generation systems, paving the way for more secure generative AI models. Warning: This paper includes model outputs that may be considered offensive.