Domain-robust vision transformer with hierarchical swin encoding for explainable low-latency driver drowsiness

Al Rafy1, Md Najmul Gony2, Md Mashfiquer Rahman3

  • 1College of Technology and Engineering, Westcliff University, Irvine, CA, 92614, USA.

Scientific Reports
|June 19, 2026
PubMed
Summary

A new lightweight Vision Transformer (ViT) model, Light-VTD, accurately detects driver drowsiness from facial images. It offers explainable heatmaps and low-latency performance on devices like Raspberry Pi 4.

Related Concept Videos