Related Experiment Video
Updated: Aug 2, 2026

10:25
Deep Learning-Based Segmentation of Cryo-Electron Tomograms
Published on: November 11, 2022
9.0K
Transformer-Based Semantic Segmentation for Extraction of Building Footprints from Very-High-Resolution Images
Jia Song1,2, A-Xing Zhu1,3, Yunqiang Zhu1
1State Key Laboratory of Resources and Environmental Information System, Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences, Beijing 100101, China.
Sensors (Basel, Switzerland)
|June 10, 2023
Summary
Vision Transformer networks enhance semantic segmentation for extracting building footprints from very-high-resolution (VHR) images. Optimizing hyperparameters like image patch size and embedding dimensions improves accuracy, outperforming traditional convolutional neural networks (CNNs).
Area of Science:
- Computer Vision
- Remote Sensing
- Deep Learning
Background:
- Semantic segmentation is crucial for object extraction in very-high-resolution (VHR) remote sensing images.
- Vision Transformer networks offer improved performance over Convolutional Neural Networks (CNNs) for semantic segmentation tasks.
Purpose of the Study:
- To investigate the impact of Vision Transformer network hyperparameters on building footprint extraction accuracy in VHR images.
- To compare the performance of Transformer-based models against CNNs for this task.
Main Methods:
- Designed and compared Transformer-based models with varying hyperparameter configurations (image patches, linear embedding dimensions, multi-head self-attention).
- Analyzed the influence of these hyperparameters on the accuracy of building footprint extraction.
Main Results:
- Smaller image patches and higher-dimension embeddings significantly improve segmentation accuracy.
- Transformer-based networks achieve higher accuracy than CNNs with comparable model sizes and training times.
- The models are scalable and can be trained on general-purpose GPUs.
Conclusions:
- Vision Transformer networks show significant potential for accurate object extraction in VHR remote sensing imagery.
- Hyperparameter tuning is essential for optimizing Transformer performance in this domain.
- These findings offer valuable insights for future remote sensing applications using deep learning.

