Related Experiment Video
Updated: Aug 5, 2026

13:26
Measuring Connectivity in the Primary Visual Pathway in Human Albinism Using Diffusion Tensor Imaging and Tractography
Published on: August 11, 2016
DiffRIS: Enhancing referring remote sensing image segmentation with pre-trained text-to-image diffusion models
Zhe Dong1, Yu-Zhe Sun1, Tian-Zhu Liu1
1School of Electronics and Information Engineering, Harbin Institute of Technology, Harbin 150001, China.
Fundamental Research
|August 1, 2026
Summary
DiffRIS enhances remote sensing image segmentation by using diffusion models to better align text descriptions with aerial images. This novel framework achieves state-of-the-art results in precise region delineation for critical applications.
Area of Science:
- Computer Science
- Remote Sensing
- Artificial Intelligence
Background:
- Referring remote sensing image segmentation (RRSIS) is crucial for applications like disaster response and urban planning.
- Current RRSIS methods struggle with aerial imagery's scale variations, diverse orientations, and semantic ambiguities.
Purpose of the Study:
- To introduce DiffRIS, a novel framework leveraging pre-trained text-to-image diffusion models for improved RRSIS.
- To enhance cross-modal alignment between natural language descriptions and remote sensing imagery.
Main Methods:
- Developed a context perception adapter (CP-adapter) for refining linguistic features via global context and object-aware reasoning.
- Introduced a progressive cross-modal reasoning decoder (PCMRD) for iterative alignment of text and visual regions.
- Utilized pre-trained diffusion models to bridge the domain gap in remote sensing vision-language tasks.
Main Results:
- DiffRIS consistently outperformed existing methods on RRSIS-D, RefSegRS, and RISBench datasets.
- Achieved new state-of-the-art performance across standard segmentation metrics.
- Demonstrated significant improvements in precise region delineation using natural language prompts.
Conclusions:
- The proposed DiffRIS framework effectively leverages diffusion models for advanced RRSIS.
- The CP-adapter and PCMRD innovations enable fine-grained semantic alignment and robust performance.
- DiffRIS establishes a new benchmark for text-guided segmentation in remote sensing imagery.
