Related Experiment Video
Updated: Mar 18, 2026

04:23
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
2.4K
Hybrid attention optimized hierarchical multiscale transformer architecture for image super-resolution
Biao Wang1,2, Rong Gao2, Ting Zhou2
1School of Information Engineering, Wuhan College, Wuhan, 430212, China.
Scientific Reports
|March 17, 2026
Summary
The new Hierarchical Multiscale Transformer Architecture (HMT) improves image super-resolution by better modeling local features and enhancing detail restoration. This transformer model achieves superior performance in generating clearer, more coherent high-resolution images.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Deep Learning
Background:
- Transformer architectures show promise in image super-resolution (SR).
- Existing models struggle with local feature modeling and feature representation, hindering fine detail restoration in high-resolution (HR) images.
Purpose of the Study:
- To address limitations in current Transformer-based SR models.
- To propose a novel architecture for improved local feature interaction and detail reconstruction.
Main Methods:
- Introduced the Hierarchical Multiscale Transformer Architecture (HMT) with Hybrid Attention.
- Developed the Fourier Mamba Synergy (FMSA) module for cross-domain feature collaboration and long-range dependency modeling.
- Incorporated Dynamic Convolutional Attention (DCA) for multi-scale feature enhancement and instance weight adaptation.
Main Results:
- HMT demonstrated superior performance, exceeding HGFormer by 0.07dB and 0.08dB in PSNR at a 4x scaling factor on BSD100 and Manga109 datasets.
- Achieved significant improvements in both objective metrics and visual quality on public datasets.
- The proposed architecture effectively captures fine-grained local feature interactions for clearer detail generation.
Conclusions:
- The HMT architecture offers enhanced capabilities for image super-resolution.
- It effectively resolves the issues of poor local feature modeling and limited representation in existing Transformer models.
- Results indicate strong potential for generating high-quality, detailed high-resolution images.
