Related Experiment Videos
Optimizing Trust and Safety Regions for Text-to-Image Generation in High-Dimensional Manifold Spaces
None:
Diffusion-based text-to-image (T2I) models such as Stable Diffusion (SD) and DALL $c.$ E 2 enable versatile image generation but raise significant safety concerns due to their ability to produce harmful or not-safe-for-work (NSFW) content (e.g., nudity). Existing safety strategies, including prompt filtering and machine unlearning, remain limited, as they are vulnerable to biased data, model openness, and adversarial prompt attacks. Achieving safe alignment during reinforcement learning (RL) fine-tuning is thus essential, yet faces two significant challenges: alignment fragility, where models easily lose control after optimization, and the safety-quality paradox, where improving safety often degrades visual quality. To address these issues, we propose S-TRPO, a Safety-constrained Trust-Region Policy Optimization framework that enables safe and reliable alignment of diffusion models (DMs) within the manifold policy space. S-TRPO introduces a dynamic safety-control mechanism that combines danger-region perception with trust-region constraints to maintain both safety and generation fidelity. Specifically, a KL-based safety region and a static risk model jointly evaluate harmful prompt risk and restrict unsafe deviations in policy updates. Furthermore, a Lagrangian dual-control scheme balances safety constraints with image-quality optimization. Extensive experiments on real-world adversarial benchmarks demonstrate that, under white-box UnlearnDiffAtk evaluation, S-TRPO with full malicious fine-tuning reduces the attack success rate by 51.7% relative to DPOK, while maintaining comparable image-text alignment quality. These results highlight the effectiveness of S-TRPO in mitigating risky behaviors and enhancing the reliability of T2I diffusion systems.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy