Related Experiment Video
Updated: Jan 8, 2026

04:48
Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
Published on: July 5, 2024
723
Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach
IEEE Transactions on Pattern Analysis and Machine Intelligence
|December 11, 2025
Summary
DiTFuse, an instruction-driven Diffusion Transformer, enhances image fusion by enabling semantic awareness and user control. This novel framework achieves robust, adaptable, and controllable fusion across various modalities without ground-truth data.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Image Processing
Background:
- Existing image fusion methods lack robustness, adaptability, and user control, especially in challenging conditions like low-light or color shifts.
- Current fusion networks are task-specific and struggle to incorporate high-level semantic understanding or user intent.
- The absence of ground-truth fused images and limited dataset sizes hinder the training of end-to-end models for complex fusion tasks.
Purpose of the Study:
- To introduce DiTFuse, a novel instruction-driven Diffusion Transformer framework for end-to-end, semantics-aware image fusion.
- To enable hierarchical and fine-grained control over fusion dynamics using natural language instructions.
- To overcome limitations of pre- and post-fusion pipelines by integrating semantic understanding directly into the fusion process.
Main Methods:
- DiTFuse jointly encodes images and natural language instructions in a shared latent space for instruction-driven fusion.
- A multi-degradation masked-image modeling strategy trains the network for cross-modal alignment and modality-invariant restoration without ground truth.
- A curated instruction dataset facilitates interactive fusion capabilities and zero-shot generalization to new fusion scenarios.
Main Results:
- DiTFuse demonstrates superior quantitative and qualitative performance on infrared-visible, multi-focus, and multi-exposure fusion benchmarks.
- The framework achieves sharper textures and improved semantic retention compared to existing methods.
- Experiments confirm DiTFuse's ability to support multi-level user control and generalize to instruction-conditioned segmentation tasks.
Conclusions:
- DiTFuse offers a unified, instruction-driven approach to image fusion, significantly advancing robustness, adaptability, and controllability.
- The model effectively integrates semantic understanding and user intent, overcoming key limitations of prior fusion techniques.
- DiTFuse presents a versatile architecture capable of handling diverse fusion tasks and enabling novel applications like instruction-conditioned segmentation.

