Related Experiment Video
Updated: Aug 13, 2025

04:23
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
1.9K
Convolutional Neural Networks or Vision Transformers: Who Will Win the Race for Action Recognitions in Visual Data?
Oumaima Moutik1, Hiba Sekkat1, Smail Tigani1
1Engineering Unit, Euromed Research Center, Euro-Mediterranean University, Fes 30030, Morocco.
Sensors (Basel, Switzerland)
|January 21, 2023
Summary
This study compares Convolutional Neural Networks (CNN) and Vision Transformers (ViT) for video action recognition. It analyzes their accuracy-complexity trade-offs to determine future trends in computer vision.
Area of Science:
- Computer Vision
- Deep Learning
- Artificial Intelligence
Background:
- Action recognition in videos is a key challenge in computer vision.
- Convolutional Neural Networks (CNN) have been foundational in deep learning for visual tasks.
- Transformers, successful in NLP, are now influencing computer vision, sparking debate on their role in action recognition.
Purpose of the Study:
- To conduct a detailed study comparing CNN and Transformer models for action recognition.
- To analyze the accuracy-complexity trade-off between these two architectures.
- To discuss the future dominance of CNN or Vision Transformers (ViT) in video action recognition.
Main Methods:
- Separate analysis of CNN and Transformer models for action recognition.
- Comparative study focusing on the accuracy-complexity trade-off.
- Performance analysis to inform conclusions.
Main Results:
- Detailed comparison of CNN and Transformer performance in action recognition tasks.
- Evaluation of the balance between model accuracy and computational complexity.
- Identification of key trends and potential future directions.
Conclusions:
- The study provides insights into the comparative strengths and weaknesses of CNN and ViT for action recognition.
- The findings will guide the choice of architecture based on specific application requirements.
- Discussion on whether CNN or Vision Transformers will ultimately lead in this field.
Related Concept Videos
Transformers with Off-Nominal Turns Ratios
189
In scenarios involving parallel transformers with disparate ratings, developing per-unit models requires accommodating off-nominal turns ratios. This situation arises when the selected base voltages are not proportional to the transformer’s voltage ratings. Consider a transformer where the rated voltages are related by the term a. If the chosen voltage bases satisfy a relationship involving term b, term c is defined as the ratio of these bases. This ratio is then substituted into the...
189
Types Of Transformers
1.0K
Transformers can provide desired voltages to a circuit by modifying the number of turns in the secondary windings.
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
1.0K
Transformers
1.1K
A device that transforms voltages from one value to another using induction is called a transformer. A transformer consists of two separate coils, or windings, wrapped around the same soft iron core. However, they are electrically insulated from each other.
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...
The iron core has a substantial relative permeability. Therefore, the magnetic field lines generated due to the current in one winding are almost entirely confined within the core, such that the same magnetic flux permeates each turn of both...
1.1K
Transformers in Distribution System
140
Transformers in distribution systems can be broadly categorized into distribution substation transformers and other distribution transformers. They are crucial for stepping down high transmission voltages to levels suitable for distribution and end-user applications.
Distribution substation transformers come in various ratings and typically use mineral oil for insulation and cooling. To prevent moisture and air from entering the oil, some transformers use an inert gas like nitrogen to fill the...
Distribution substation transformers come in various ratings and typically use mineral oil for insulation and cooling. To prevent moisture and air from entering the oil, some transformers use an inert gas like nitrogen to fill the...
140
Vision
54.7K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
54.7K
Convolution: Math, Graphics, and Discrete Signals
334
In any LTI (Linear Time-Invariant) system, the convolution of two signals is denoted using a convolution operator, assuming all initial conditions are zero. The convolution integral can be divided into two parts: the zero-input or natural response and the zero-state or forced response, with t0 indicating the initial time.
To simplify the convolution integral, it is assumed that both the input signal and impulse response are zero for negative time values. The graphical convolution process...
To simplify the convolution integral, it is assumed that both the input signal and impulse response are zero for negative time values. The graphical convolution process...
334

