Related Experiment Videos
A multi-scale feature fusion gaze estimation model based on convolutional neural network and vision transformer
Peng Wang1, Xuena Wang2, Shuo Yuan3
1School of Electrical and Information Engineering, Changzhou Institute of Technology, Changzhou, 213032, China. wangp@czu.cn.
Scientific Reports
|June 1, 2026
Summary
This study introduces CAF-ViT, a novel multi-scale fusion model for accurate gaze estimation. It significantly improves performance in unconstrained environments by enhancing feature fusion and reducing feature loss.
Area of Science:
- Computer Vision
- Machine Learning
- Artificial Intelligence
Background:
- Gaze estimation in unconstrained environments faces challenges with ineffective feature fusion and feature loss.
- Existing methods struggle to effectively integrate multi-scale features for robust gaze prediction.
Purpose of the Study:
- To propose a novel multi-scale feature fusion model, CAF-ViT (Cross-Attention Fusion Vision Transformer), to address limitations in current gaze estimation techniques.
- To enhance feature representation and fusion by integrating coarse- and fine-grained details from multi-scale inputs.
Main Methods:
- Inputting multi-scale face images into ResNet-18 for feature extraction at different granularities.
- Employing learnable Class Tokens per scale, followed by self-attention for local and global feature aggregation.
- Implementing cross-attention between different scale token sequences for deep bidirectional feature fusion and an additional attention layer for refinement.
Main Results:
- Achieved estimation errors of [Formula: see text] on MPIIFaceGaze, an 8.6% improvement over the baseline.
- Attained errors of [Formula: see text] on EyeDiap and [Formula: see text] on Gaze360 datasets.
- Ablation studies confirmed the effectiveness of the multi-scale fusion strategy and the improved attention mechanism.
Conclusions:
- The proposed CAF-ViT model effectively overcomes feature fusion and loss issues in unconstrained gaze estimation.
- The multi-scale fusion and refined attention mechanisms contribute to state-of-the-art performance across diverse datasets.
- CAF-ViT offers a promising advancement for accurate and robust gaze direction prediction.