Related Experiment Video
Updated: May 11, 2026

04:48
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
2.7K
A transformer-based approach empowered by a self-attention technique for semantic segmentation in remote sensing
Wadii Boulila1,2, Hamza Ghandorh3, Sharjeel Masood4
1Robotics and Internet-of-Things Laboratory, Prince Sultan University, Riyadh 12435, Saudi Arabia.
Heliyon
|April 26, 2024
Summary
This study introduces a novel deep learning method for remote sensing image segmentation, combining convolutional and transformer architectures to accurately identify fine details and small objects with reduced computational needs.
Area of Science:
- Computer Vision
- Remote Sensing Image Analysis
- Deep Learning Architectures
Background:
- Semantic segmentation of remote sensing (RS) images is vital for land cover classification and scene understanding.
- Existing deep learning methods struggle with fine details and high computational costs in RS image analysis.
- Processing fine details in high-resolution RS images while managing computational demands remains a challenge.
Purpose of the Study:
- To develop a novel approach for semantic segmentation of RS images that effectively processes fine details and small objects.
- To address the computational demands associated with high-resolution RS image analysis.
- To improve the accuracy of semantic segmentation in remote sensing applications.
Main Methods:
- A hybrid deep learning architecture combining convolutional layers for fine-grained features and transformer blocks for contextual information.
- Utilizing convolutional layers with a low receptive field to generate detailed feature maps for small objects.
- Employing transformer blocks to capture global contextual information, reducing the need for extensive downsampling and enabling full-resolution feature processing.
Main Results:
- The proposed method achieved a mean Dice score of 80.41%, outperforming UNet (78.57%), FCN (74.57%), PSP Net (73.45%), and CvT (62.97%).
- Demonstrated effectiveness in generating both local and contextual features crucial for accurate segmentation.
- The approach requires less extensive datasets compared to purely transformer-based networks.
Conclusions:
- The novel hybrid convolutional-transformer approach significantly enhances semantic segmentation accuracy for remote sensing images, particularly for small objects.
- This method offers a computationally efficient solution for processing high-resolution remote sensing data.
- The findings suggest a promising direction for improving scene understanding and land cover classification in remote sensing.

