Related Experiment Videos
DSES-Diff: A Dynamic-Spectral and Efficient-Spatial Diffusion-Based Foundation Model for Scalable Hyperspectral and
Abstract:
Hyperspectral image (HSI) and multispectral image (MSI) fusion is a critical technique for generating imagery with high spatial and spectral resolution. However, developing robust fusion models is fundamentally challenged by the severe scarcity of large-scale, precisely paired HSI-MSI datasets, which are prohibitively difficult to acquire. This data limitation constrains the performance and generalization of existing data-driven methods. Furthermore, effectively leveraging the complementary strengths - the rich spectral prior of HSIs and the detailed spatial prior of MSIs - remains a complex problem. To overcome this limitation, we propose DSES-Diff, a foundation model for HSI-MSI fusion that eliminates the dependence on large-scale paired data. Our approach separately trains a dynamic-spectral network and an efficient-spatial network using extensive unpaired high-resolution spectral and spatial datasets to extract strong modality-specific priors. Specifically, we introduce AdaSpecNet, a spectral network built upon the Diffusion Transformer (DiT) architecture, which treats each individual spectral band as a sequence token. By leveraging positional padding and attention masks, AdaSpec-Net can flexibly accommodate scalable spectral inputs with band numbers in a unified manner. Furthermore, we propose MoESpatNet, a spatial network that incorporates a Mixture-of-Experts (MoE) module to enhance the network's ability to model diverse spatial distributions. The MoE design enables sparse expert activation during inference, improving fusion quality and computational efficiency. Extensive experiments across various datasets demonstrate that our method achieves superior fusion performance under diverse conditions, showcasing its strong potential to serve as a versatile foundation for scalable HSI and MSI fusion.