Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Jun 21, 2026

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
04:23

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

Published on: April 21, 2023

Domain-robust vision transformer with hierarchical swin encoding for explainable low-latency driver drowsiness

Al Rafy1, Md Najmul Gony2, Md Mashfiquer Rahman3

  • 1College of Technology and Engineering, Westcliff University, Irvine, CA, 92614, USA.

Scientific Reports
|June 19, 2026
PubMed
Summary

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Explainable detection of goldenhar syndrome using an OD-Mamba backbone for rare craniofacial disorder diagnosis.

Scientific reports·2026
Same author

Temporal and Regional Circular RNA profiling in a Tauopathy Mouse Model: Implications for Tau Pathology and Neurodegeneration.

bioRxiv : the preprint server for biology·2026
Same author

Emotionally expressive facial animation driven by EmotionBERT embeddings.

Scientific reports·2026
Same author

Deep ensemble of multi-head attention CNNs for histopathological image-based of lung and colon cancer diagnosis.

Digital health·2026
Same author

Explainable clustering and ranking of student satisfaction in online learning: a hybrid FAHP-TOPSIS and SHAP-LIME approach.

Scientific reports·2026
Same author

Explainable AI-driven hybrid deep learning framework for accurate skin cancer diagnosis.

Digital health·2026

A new lightweight Vision Transformer (ViT) model, Light-VTD, accurately detects driver drowsiness from facial images. It offers explainable heatmaps and low-latency performance on devices like Raspberry Pi 4.

Area of Science:

  • Computer Vision
  • Machine Learning
  • Artificial Intelligence

Background:

  • Driver drowsiness is a major cause of road accidents globally.
  • Existing fatigue detection systems struggle with generalization, transparency, and low-power deployment.
  • There is a need for accurate, interpretable, and efficient fatigue monitoring solutions.

Purpose of the Study:

  • To introduce Light-VTD, a novel lightweight Vision Transformer (ViT) architecture for accurate and explainable driver fatigue detection.
  • To improve upon existing methods by enhancing local feature extraction, global context capture, and spatial consistency.
  • To enable real-time, low-power fatigue monitoring with visual explanations.

Main Methods:

  • Developed Light-VTD, incorporating fused-MBConv blocks, a position-aware token mixer, and multi-path token fusion.
Keywords:
Drowsiness detectionEmbedded systemsExplainable AI (XAI)Token fusionVision transformer (ViT)

More Related Videos

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
04:48

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

Published on: July 5, 2024

Related Experiment Videos

Last Updated: Jun 21, 2026

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
04:23

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

Published on: April 21, 2023

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
04:48

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

Published on: July 5, 2024

  • Adapted Gradient-weighted Class Activation Mapping (Grad-CAM) for interpretable attention heatmaps.
  • Trained and evaluated the model on a large-scale dataset (167,000+ images) from four diverse public datasets.
  • Main Results:

    • Achieved 99.2% intra-dataset accuracy and over 93% cross-dataset transfer accuracy.
    • Outperformed existing ViT baselines by 2-4% in accuracy.
    • Demonstrated low-latency inference (<120 ms) suitable for low-power devices (Raspberry Pi 4).

    Conclusions:

    • Light-VTD provides a robust, interpretable, and deployable solution for driver fatigue detection.
    • The model's explainability through Grad-CAM aids in understanding fatigue indicators.
    • Potential applications include Advanced Driver Assistance Systems (ADAS), workplace safety, and portable diagnostics.