Related Experiment Videos
A unified multi-modal latent diffusion framework with modality-dropout training
S Remya1, Manu J Pillai2, Laveena Herman2
1Department of CSE, Amrita School of Computing, Amrita Vishwa Vidyapeetham, Amritapuri, Kollam, Kerala, India.
Introduction:
Multi-modal conditioning in latent diffusion models-combining text, structural, and spatial guidance signals-substantially improves controllable image synthesis, yet two limitations persist across most existing frameworks. First, conditioning modalities are fused using fixed architectural weights that remain constant regardless of the input, preventing any input-adaptive rebalancing of guidance. Second, most existing multi-modal approaches are designed under the assumption that all conditioning signals will be present at inference time, leaving the model without any way to degrade gracefully when some modalities are unavailable. This paper introduces a unified latent diffusion framework that addresses both limitations simultaneously.
Methods:
Our architecture integrates three complementary conditioning modalities-semantic text embeddings via CLIP, structural boundary maps derived from Canny edge detection, and regional semantic layout from segmentation masks-as co-equal inputs into a single U-Net denoising backbone. The framework has two core components: Softmax-Normalized Adaptive Guidance Fusion (SNAGF), which replaces fixed fusion weights with three learnable, softmax-normalized modality importance scores optimised end-to-end alongside the diffusion objective; and Modality-Dropout Training (MDT), a structured regularisation strategy that randomly zeroes individual guidance representations with probability p drop = 0.3 during training, training the model to produce coherent outputs under any available subset of conditioning inputs.
Results:
Evaluated on CIFAR-10, the full SNAGF+MDT framework achieves SSIM 0.93, FID 17.2, PSNR 34.1 dB, CLIP Score 0.88, LPIPS 0.09, and IS 10.4. MDT alone yields an average 15.5% FID improvement across partial-conditioning inference configurations, and learned SNAGF weights converge to λtext = 0.370, λedge = 0.363, λseg = 0.267, demonstrating stable convergence from epoch 30 onward. Results are averaged over three random seeds.
Discussion:
Taken together, they indicate that adaptive modality weighting outperforms fixed conditioning strategies and that modality-dropout training provides an effective partial-conditioning robustness mechanism. All findings are established at 32 × 32 resolution on a single benchmark with algorithmically constructed conditioning signals; validation on higher-resolution datasets with naturally paired text and structural annotation remains necessary before these conclusions can be generalised.
Related Concept Videos
Multicompartment Models: Overview
These models offer a more comprehensive representation of drug behavior in the body than one-compartment models. They accommodate the complexity of drug distribution,...
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...