Related Experiment Video
Updated: Jul 30, 2026

04:23
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
2.2K
A novel deep transformer based CvT model for sign language recognition in visual communication
Scientific Reports
|December 21, 2025
Summary
A new Convolutional Vision Transformer (CvT) model significantly improves sign language recognition (SLR) accuracy to 99%. This AI advancement enhances communication accessibility for deaf and hard-of-hearing communities by overcoming limitations of traditional methods.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Sign language recognition (SLR) is vital for deaf and hard-of-hearing communities.
- Traditional AI and deep learning models face challenges with complex gestures, lighting, and occlusions.
- Vision Transformers (ViTs) offer advanced feature extraction via self-attention, modeling long-range dependencies effectively.
Purpose of the Study:
- To develop an advanced AI model for accurate and robust sign language recognition.
- To integrate convolutional and transformer mechanisms for optimized feature extraction.
- To evaluate the proposed model's performance against existing state-of-the-art methods.
Main Methods:
- A Convolutional Vision Transformer (CvT) model was proposed, combining hierarchical convolutional tokenization with transformer attention.
- The CvT model was trained and evaluated on sign language digits and alphabet/symbol datasets.
- Performance was benchmarked against traditional Convolutional Neural Networks (CNNs) and transformer-based models like BeIT.
Main Results:
- The proposed CvT model achieved a 99% accuracy on both evaluated datasets.
- CvT demonstrated superior performance compared to baseline CNN and BeIT models.
- The model effectively reduced misclassifications and improved predictive confidence and generalization.
Conclusions:
- The CvT model represents a significant advancement in automated sign language recognition.
- This AI approach enhances the accuracy and reliability of SLR systems.
- The findings support the potential of CvT for more inclusive communication technologies.