Related Experiment Videos
RS-CARES: Context-Aware Cross-Modal Alignment with Semantic Spatial Prior for Referring Remote Sensing Image
Hui Xiong1, Wen Luo1, Bing He1
1School of Computer Science, Chengdu University of Information Technology, Chengdu 610225, China.
Abstract:
Remote sensing referring image segmentation faces critical challenges including arbitrary target rotation, drastic scale variation, cluttered complex backgrounds, and large visual-language semantic gaps. Existing mainstream segmentation models adopt fixed-receptive-field backbones, coarse unidirectional cross-modal interaction and static learnable object queries, which easily cause small-object missed detection, blurred boundary segmentation and fragmented predictions. This paper proposes a context-aware referring expression segmentation model for remote sensing to tackle the above limitations. We construct a dual-stream feature extraction backbone with InternImage and CLIP text encoder, and design a semantic prior localization map module to generate spatial heatmaps for spatial inductive bias and improve small-object localization recall. A cross-modal context aggregator performs multi-scale bidirectional visual-text alignment, while a dynamic query initialization strategy and a language-guided Transformer decoder progressively improve target localization and mask refinement. Experiments are conducted on two standard RRSIS benchmarks, RefSegRS and RRSIS-D. On RefSegRS, the proposed model achieves an mIoU of 70.43% and an oIoU of 77.34%, and obtains significant improvements on small vehicles, buildings and slender road markings. On the larger-scale RRSIS-D dataset, our model achieves an mIoU of 63.71% and an oIoU of 73.68%. The proposed model provides a solution for accurate multi-scale target segmentation guided by natural language descriptions in complex remote sensing scenes.