Related Experiment Video
Updated: Jan 14, 2026

Experimental Paradigm for Measuring the Effects of Self-distancing in Young Children
Published on: March 1, 2019
The Safety Illusion? Testing the Boundaries of Concept Removal in Diffusion Models
None:
Text-to-image diffusion models are capable of producing high-quality images from textual descriptions; however, they present notable security concerns. These include the potential for generating Not-Safe-For-Work (NSFW) content, replicating artists' styles without authorization, or creating deepfakes. Recent advancements have proposed concept erasure techniques to eliminate sensitive concepts from these models, aiming to mitigate the generation of undesirable content. Nevertheless, the robustness of these techniques against a wide range of adversarial inputs has not been comprehensively investigated. To address this challenge, a novel two-stage optimization attack framework based on adversarial perturbations, referred to as Concept Embedding Adversary (CEA), was proposed in the present study. By leveraging the cross-modal alignment priors of the CLIP model, CEA iteratively adjusts adversarial embedding vectors to approximate the semantic expression of specific target concepts. This process enables the construction of deceptive adversarial prompts that exploit diffusion models, compelling them to regenerate previously erased concepts. The performance of concept erasure methods was evaluated, specifically when dealing with diversified adversarial prompts targeting erased concepts, such as NSFW content, artistic styles, and objects. Extensive experimental results demonstrate that existing concept erasure methods are unable to completely eliminate target concepts. In contrast, the proposed CEA framework exploits residual vulnerabilities within the generative latent space through a two-stage optimization process. By achieving precise cross-modal alignment, CEA attains significantly higher ASR in regenerating erased concepts.
More Related Videos
06:42Continuous Theta Burst Stimulation of the Posterior Medial Frontal Cortex to Experimentally Reduce Ideological Threat Responses
Published on: September 28, 2018
07:34Perceptual and Category Processing of the Uncanny Valley Hypothesis' Dimension of Human Likeness: Some Methodological Issues
Published on: June 3, 2013
Related Concept Videos
Theories of Dissolution: Diffusion Layer Model
This process starts with a thin layer, saturated with the drug, forming at the interface between the solid and liquid. The solute then diffuses from this layer into the main solution. The Noyes-Whitney equation suggests that the rate of dissolution relies on the diffusion...
Theories of Dissolution: The Danckwerts' Model and Interfacial Barrier Model
Deindividuation
Magical Thinking
Causes of Similarity-Dissimilarity Effect
The Scientific Method