Related Experiment Video
Updated: Aug 12, 2026

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
Extending the scale generalization of the Vision Transformer without fine-tuning
Kai Jiang1, Peng Peng1, Youzao Lian2
1Department of Control Science and Engineering, College of Electronic and Information Engineering, Tongji University, No. 4800, Caoan Highway, Shanghai, 201804, Shanghai, China.
Vision Transformers (ViTs) struggle with different image resolutions. The new Multi-Scale Vision Transformer (MSViT) improves resolution generalization by addressing spectral drift and positional awareness issues.
Area of Science:
- Computer Vision
- Deep Learning
- Artificial Intelligence
Background:
- The "train low, deploy high" paradigm is practical but challenging for Vision Transformers (ViTs).
- ViTs exhibit poor zero-shot generalization to unseen resolutions, unlike convolutional networks.
- Key issues include intra-patch spectral drift and inter-patch positional awareness collapse during resizing.
Purpose of the Study:
- To develop a resolution-flexible Vision Transformer that maintains high-fidelity inference across various spatial resolutions.
- To address the limitations of ViTs in handling different input scales.
- To improve zero-shot generalization capabilities of ViTs.
Main Methods:
- Proposed the Multi-Scale Vision Transformer (MSViT).
- Integrated Spectral-Constrained Convolution for adaptive frequency-weighted patch embedding.
- Employed Horizontal-Vertical Separable Attention for a full-span receptive field.
- Utilized Reparameterized Convolutional Position Embedding for boundary-aware spatial bias.
Main Results:
- MSViT demonstrated remarkable robustness across resolutions from 128x128 to 640x640.
- Consistent and stable accuracy was maintained when trained solely at 224x224.
- The model effectively mitigated intra-patch spectral drift and inter-patch positional awareness collapse.
Conclusions:
- Explicit modeling of spectral stability and spatial structure is crucial for resolution-flexible ViTs.
- MSViT offers a viable solution for the "train low, deploy high" paradigm in Vision Transformers.
- The proposed methods enhance ViT performance and generalization across diverse input scales.
Related Concept Videos
Transformers with Off-Nominal Turns Ratios
Depth Perception and Spatial Vision
Transformations of Functions III

