Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Videos

Optimizing Trust and Safety Regions for Text-to-Image Generation in High-Dimensional Manifold Spaces.

Xiang Yang, Xiaohui Li, Yuke Wang

    IEEE Transactions on Pattern Analysis and Machine Intelligence
    |June 2, 2026
    PubMed
    Summary

    We introduce S-TRPO, a novel framework for safe text-to-image generation. S-TRPO effectively mitigates harmful content generation in diffusion models while preserving image quality, enhancing AI safety.

    Related Concept Videos

    Improving Translational Accuracy02:07

    Improving Translational Accuracy

    Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
    Improving Translational Accuracy02:07

    Improving Translational Accuracy

    Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...

    You might also read

    Related Articles

    Articles linked to this work by shared authors, journal, and citation graph.

    Sort by
    Same author

    Interpretable agentic AI system with localized reasoning for radiology.

    NPJ digital medicine·2026
    Same author

    SymBOL: A General-Purpose Symbolic Learner for Scientific Discovery Using Bayesian Optimization-Enhanced Large Language Models.

    IEEE transactions on pattern analysis and machine intelligence·2026
    Same author

    LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology.

    ACM transactions on computing for healthcare·2026
    Same author

    Influencing factors of functional exercise adherence in stroke survivors: a cross-sectional study based on structural equation modeling.

    Scientific reports·2026
    Same author

    Contrastive multi-view representation learning for multi-camera plant phenotyping: A cotton field study.

    Plant phenomics (Washington, D.C.)·2026
    Same author

    MCF-YOLO: Consistency-Guided Cross-Modal Attention for Small-Object RGB-IR Detection.

    Sensors (Basel, Switzerland)·2026

    Area of Science:

    • Artificial Intelligence
    • Computer Vision
    • Machine Learning

    Background:

    • Diffusion models (e.g., Stable Diffusion, DALL-E 2) excel at text-to-image generation but pose safety risks due to harmful content.
    • Existing safety measures like prompt filtering and unlearning are insufficient against data bias and adversarial attacks.
    • Reinforcement learning (RL) fine-tuning for safety faces challenges like alignment fragility and the safety-quality paradox.

    Purpose of the Study:

    • To develop a robust framework for safe alignment of diffusion models.
    • To address the limitations of current safety strategies in text-to-image generation.
    • To improve the reliability and safety of diffusion models without compromising visual quality.

    Main Methods:

    • Proposed S-TRPO (Safety-constrained Trust-Region Policy Optimization) framework for diffusion models.

    Related Experiment Videos

  • Implemented a dynamic safety-control mechanism using danger-region perception and trust-region constraints.
  • Utilized a KL-based safety region and a static risk model for harmful prompt evaluation.
  • Employed a Lagrangian dual-control scheme to balance safety and image quality.
  • Main Results:

    • S-TRPO significantly reduces the attack success rate by 51.7% compared to DPOK under white-box UnlearnDiffAtk evaluation.
    • Maintained comparable image-text alignment quality alongside enhanced safety.
    • Demonstrated effectiveness in mitigating risky behaviors in text-to-image diffusion systems.

    Conclusions:

    • S-TRPO provides a reliable method for safe alignment of diffusion models.
    • The framework effectively balances safety constraints with the optimization of image generation quality.
    • S-TRPO enhances the overall safety and trustworthiness of text-to-image generation technologies.