Related Experiment Video
Updated: Oct 1, 2026

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
Published on: July 5, 2024
ALSRFormer: An adaptive transformer with dynamic window attention and multi-scale deformable feed-forward network for
Ziqi Jia1, Jiawei Zhang1, Dongen Guo1
1School of Artificial Intelligence, Nanyang Institute of Technology, Nanyang, 473004, China.
Abstract:
High-resolution remote sensing images present considerable challenges for semantic segmentation due to their complex object structures and extensive spatial distribution. Effective segmentation requires capturing fine-grained local details while simultaneously modeling long-range dependencies. Convolutional Neural Networks (CNNs) are well-suited for local detail extraction, whereas Vision Transformers (ViTs) excel at global dependency modeling, yet both exhibit inherent limitations. Although CNN-Transformer hybrid approaches have improved global feature representation, their reliance on fixed window partitioning, static feature fusion, and rigid convolutional kernels constrains performance on irregular objects and complex textures. To address these limitations, we propose the Adaptive Long Short Range Transformer (ALSRFormer) built upon the ConvNeXt backbone, forming the ConvALSR-Net model. Specifically, we design a Dynamic Window Multi-head Self-Attention (DynamicWMSA) mechanism that adaptively adjusts local attention windows according to image content. We further introduce a Multi-Scale Deformable Directional FFN (MSDD-FFN) to enhance the representation of complex textures and boundaries via multi-scale deformable convolutions. Additionally, Inception Depthwise Convolution (IDConv) and an AdaptiveFusion module are integrated to refine shallow features and improve global information fusion. Experiments on the LoveDA, Vaihingen, and Potsdam datasets show that ConvALSR-Net consistently outperforms existing methods, with mIoU improvements ranging from 0.6 to 1.1 percentage points over competitive baselines while using fewer parameters (65.15M) and lower FLOPs (70.30G) than the original ConvLSR-Net. The gains are more evident on the LoveDA dataset, where complex multi-scale scenes amplify the benefit of content-adaptive modeling.