Related Experiment Video
Updated: Mar 13, 2026

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
Published on: July 5, 2024
CSAFusion: a convolutional neural network (CNN)-based and Swin Transformer network for multi-modal medical image
Liyuan Zhang1,2, Jiachen Zheng1,2, Xiongfeng Tang3
1School of Computer Science and Technology, Changchun University of Science and Technology, Changchun, China.
Background:
Single-modality medical imaging provides limited clinical information, whereas multi-modal fusion aggregates complementary structural and functional data to improve diagnosis. Multi-scale decomposition-based fusion methods have attracted attention for their ability to separate and process image features at different resolutions, but they still struggle to preserve spatial fidelity: fine structural details are often lost, particularly at tissue boundaries and in texture-rich regions. This study aimed to develop a multi-modal medical image fusion method that better preserves spatial fidelity and fine structural details.
Methods:
We propose an unsupervised fusion framework termed CSAFusion, which adopts a U-Net backbone with adaptive convolutions (ACs) in the encoder and Swin Transformer modules in the decoder. The network introduces three main innovations: (I) a dual fusion architecture that integrates complementary information at both encoder and decoder levels; (II) AC kernels for content-aware, dynamic feature extraction; and (III) a composite loss function incorporating structural similarity, regional mutual information, and contrast preservation to jointly optimize structural integrity and perceptual quality. Quantitative performance was evaluated using seven metrics: structural similarity index (SSIM), spatial frequency (SF), average gradient (AG), edge preservation index (EPI), feature mutual information (FMI), visual information fidelity (VIF), and non-reference quality index for enhancement (QNCIE). Statistical significance of differences among fusion methods was assessed using repeated-measures analyses, with P<0.05 considered statistically significant.
Results:
Extensive experiments on computed tomography-magnetic resonance imaging (CT-MRI), T1- and T2-weighted magnetic resonance imaging (MRI-T1-T2), and MRI-single-photon emission computed tomography (MRI-SPECT) fusion tasks demonstrated that CSAFusion consistently outperformed comparison methods, particularly in preserving CT bone structures, MRI soft-tissue textures, and SPECT functional information. The proposed method achieved a QNCIE score of 0.9304, corresponding to an approximate 3% improvement over existing approaches. For MRI-SPECT fusion, CSAFusion yielded an SSIM of 0.8882±0.0824 and an EPI of 0.9002±0.0187. Across all three modality pairs, CSAFusion obtained the highest or near-highest values for all seven quantitative metrics, and its advantages over competing methods were statistically significant.
Conclusions:
CSAFusion enhances both local and global feature extraction, enabling more faithful preservation of critical diagnostic information in multi-modal medical image fusion. The superior quantitative performance and robust cross-modality generalization highlight CSAFusion as a promising tool for clinical applications.
