Related Experiment Videos

A multi-scale feature fusion gaze estimation model based on convolutional neural network and vision transformer

Peng Wang1, Xuena Wang2, Shuo Yuan3

  • 1School of Electrical and Information Engineering, Changzhou Institute of Technology, Changzhou, 213032, China. wangp@czu.cn.

Scientific Reports
|June 1, 2026
PubMed
Summary

This study introduces CAF-ViT, a novel multi-scale fusion model for accurate gaze estimation. It significantly improves performance in unconstrained environments by enhancing feature fusion and reducing feature loss.

Related Concept Videos