Related Experiment Video
Updated: Jun 21, 2026

04:23
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
Domain-robust vision transformer with hierarchical swin encoding for explainable low-latency driver drowsiness
Al Rafy1, Md Najmul Gony2, Md Mashfiquer Rahman3
1College of Technology and Engineering, Westcliff University, Irvine, CA, 92614, USA.
Scientific Reports
|June 19, 2026
Summary
A new lightweight Vision Transformer (ViT) model, Light-VTD, accurately detects driver drowsiness from facial images. It offers explainable heatmaps and low-latency performance on devices like Raspberry Pi 4.
Area of Science:
- Computer Vision
- Machine Learning
- Artificial Intelligence
Background:
- Driver drowsiness is a major cause of road accidents globally.
- Existing fatigue detection systems struggle with generalization, transparency, and low-power deployment.
- There is a need for accurate, interpretable, and efficient fatigue monitoring solutions.
Purpose of the Study:
- To introduce Light-VTD, a novel lightweight Vision Transformer (ViT) architecture for accurate and explainable driver fatigue detection.
- To improve upon existing methods by enhancing local feature extraction, global context capture, and spatial consistency.
- To enable real-time, low-power fatigue monitoring with visual explanations.
Main Methods:
- Developed Light-VTD, incorporating fused-MBConv blocks, a position-aware token mixer, and multi-path token fusion.
- Adapted Gradient-weighted Class Activation Mapping (Grad-CAM) for interpretable attention heatmaps.
- Trained and evaluated the model on a large-scale dataset (167,000+ images) from four diverse public datasets.
Main Results:
- Achieved 99.2% intra-dataset accuracy and over 93% cross-dataset transfer accuracy.
- Outperformed existing ViT baselines by 2-4% in accuracy.
- Demonstrated low-latency inference (<120 ms) suitable for low-power devices (Raspberry Pi 4).
Conclusions:
- Light-VTD provides a robust, interpretable, and deployable solution for driver fatigue detection.
- The model's explainability through Grad-CAM aids in understanding fatigue indicators.
- Potential applications include Advanced Driver Assistance Systems (ADAS), workplace safety, and portable diagnostics.
