Related Experiment Video
Updated: Jun 25, 2026

12:30
A Comprehensive Protocol for Manual Segmentation of the Medial Temporal Lobe Structures
Published on: July 2, 2014
Semantic consistency-aware pseudo-temporal framework for multimodal remote sensing image segmentation
Yujia Sun1, Yuejiang Li1, Weisheng Dong1
1School of Artificial Intelligence, Xidian University, Xi'an, 710071, China.
Summary
This study introduces a novel framework for multimodal remote sensing semantic segmentation in ultra-high-resolution imagery. The approach improves spatial context and cross-modal consistency for better land-cover recognition.
Area of Science:
- Geospatial analysis
- Computer vision
- Remote sensing
Background:
- Multimodal remote sensing semantic segmentation is crucial for land-cover recognition but struggles with ultra-high-resolution imagery.
- Existing methods suffer from spatial context fragmentation and inconsistent multimodal fusion, limiting generalization.
Purpose of the Study:
- To propose a semantic consistency-aware pseudo-temporal multimodal segmentation framework for ultra-high-resolution remote sensing imagery.
- To address challenges in capturing long-range dependencies and ensuring cross-modal semantic consistency.
Main Methods:
- Developed a framework utilizing the Segment Anything Model (SAM).
- Introduced a pseudo-temporal input construction strategy using random walks and image transformations.
- Implemented a cross-modal temporal interaction module with pyramid fusion and cross-frame attention.
- Employed a prompt-guided decoder with semantic similarity constraints.
Main Results:
- The proposed framework effectively models cross-region contextual dependencies and cross-modal semantic complementarity.
- Achieved significant performance improvements over existing methods on ISPRS Vaihingen and Potsdam datasets.
- Demonstrated enhanced class separability and structural consistency in segmentation results.
Conclusions:
- The semantic consistency-aware pseudo-temporal framework offers a robust solution for multimodal remote sensing segmentation.
- The method shows strong generalization capabilities in complex environments and ultra-high-resolution imagery.
- Validated effectiveness for fine-grained land-cover recognition and geospatial analysis.

