Related Experiment Video
Updated: Jan 16, 2026

04:23
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
2.3K
A Comparative Survey of Vision Transformers for Feature Extraction in Texture Analysis
Leonardo Scabini1, Andre Sacilotti2, Kallil M Zielinski1
1São Carlos Institute of Physics, University of São Paulo, São Carlos 13560-970, SP, Brazil.
Journal of Imaging
|September 26, 2025
Summary
Vision Transformers (ViTs) show strong potential for texture recognition, outperforming Convolutional Neural Networks (CNNs) and traditional methods in accuracy and efficiency on GPUs. BeiTv2-B/16 achieved the highest accuracy.
Area of Science:
- Computer Vision
- Machine Learning
- Pattern Recognition
Background:
- Texture is a key visual feature in image analysis and pattern recognition.
- Convolutional Neural Networks (CNNs) are established methods for texture analysis.
- Vision Transformers (ViTs) show promise in broader visual recognition tasks.
Purpose of the Study:
- Investigate the effectiveness of Vision Transformers (ViTs) for texture recognition.
- Analyze the capabilities and limitations of various ViT architectures as feature extractors.
- Compare ViT performance against CNN-based and hand-engineered approaches.
Main Methods:
- Evaluated 25 ViT variants as feature extractors for texture analysis.
- Compared accuracy and computational efficiency of ViTs, CNNs (ResNet50), and hand-engineered methods.
- Utilized in-the-wild texture datasets and strong pre-training strategies.
Main Results:
- ViTs generally outperformed CNNs and hand-engineered models in texture recognition accuracy.
- BeiTv2-B/16 achieved the highest average accuracy (85.7%), followed by ViT-B/16-DINO (84.1%) and Swin-B (80.8%).
- Despite higher FLOPs/parameters, ViTs like ViT-B and BeiT(v2) offered faster GPU feature extraction than ResNet50.
Conclusions:
- ViTs are a powerful and efficient tool for texture analysis, surpassing traditional methods.
- Strong pre-training enhances ViT performance, especially on diverse, real-world texture datasets.
- Future research should focus on ViT efficiency improvements and domain-specific adaptations for texture recognition.
Related Concept Videos
Continuous -time Fourier Transform
831
The Fourier series is instrumental in representing periodic functions, offering a powerful method to decompose such functions into a sum of sinusoids. This technique, however, necessitates modification when applied to nonperiodic functions. Consider a pulse-train waveform consisting of a series of rectangular pulses. When these pulses have a finite period, they can be accurately represented by a Fourier series. Yet, as the period approaches infinity, resulting in a single, isolated pulse, the...
831
Vision
59.4K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
59.4K
Visual System
1.7K
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
1.7K
Depth Perception and Spatial Vision
1.8K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
1.8K

