Related Experiment Video
Updated: Jul 20, 2025

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
571
Self-Enhanced Mixed Attention Network for Three-Modal Images Few-Shot Semantic Segmentation
Kechen Song1, Yiming Zhang1, Yanqi Bao2
1School of Mechanical Engineering & Automation, Northeastern University, Shenyang 110819, China.
Sensors (Basel, Switzerland)
|July 29, 2023
Summary
This study introduces a novel three-modal (Visible-Depth-Thermal) image approach for few-shot semantic segmentation, improving performance in low-light conditions. The Self-Enhanced Mixed Attention Network (SEMANet) achieves state-of-the-art results with limited annotated data.
Area of Science:
- Computer Vision
- Machine Learning
- Image Processing
Background:
- Image segmentation is crucial but struggles with poor illumination.
- Multi-modal imaging and few-shot learning offer solutions for data scarcity and challenging conditions.
- Existing methods lack robustness in extreme low-light scenarios.
Purpose of the Study:
- To propose a novel few-shot semantic segmentation method using Visible-Depth-Thermal (three-modal) images.
- To address the performance degradation of image segmentation models in low-illumination environments.
- To introduce a new dataset and a robust network architecture for improved few-shot segmentation.
Main Methods:
- Developed a Visible-Depth-Thermal (three-modal) image dataset (VDT-2048-5i) for few-shot semantic segmentation.
- Proposed the Self-Enhanced Mixed Attention Network (SEMANet) incorporating Self-Enhanced (SE) and Mixed Attention (MA) modules.
- Utilized homogeneous and complementary information from three-modal images for feature fusion and enhancement.
Main Results:
- SEMANet achieved state-of-the-art performance, improving mean Intersection over Union (mIoU) by 3.8% (1-shot) and 3.3% (5-shot).
- The proposed method demonstrates superior performance compared to existing advanced methods in few-shot segmentation tasks.
- The SE module enhances foreground feature discrimination, while the MA module effectively fuses multi-modal features.
Conclusions:
- The three-modal Visible-Depth-Thermal approach combined with SEMANet significantly advances few-shot semantic segmentation, especially in challenging low-light conditions.
- The novel dataset and network architecture provide a strong foundation for future research in robust image segmentation.
- Future work will focus on enhancing feature representation for greater robustness and reducing computational costs.

