Related Experiment Video
Updated: Jun 8, 2025

07:09
Gaze in Action: Head-mounted Eye Tracking of Children's Dynamic Visual Attention During Naturalistic Behavior
Published on: November 14, 2018
10.6K
Spatio-Temporal Attention and Gaussian Processes for Personalized Video Gaze Estimation.
Swati Jindal1, Mohit Yadav2, Roberto Manduchi1
1University of California Santa Cruz.
Summary
This study introduces a novel deep learning model for video gaze estimation, improving accuracy with a spatial attention mechanism and personalization. The method enhances understanding of human attention and behavior from facial videos.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Human-Computer Interaction
Background:
- Gaze analysis is crucial for understanding human behavior and attention.
- Estimating gaze direction from facial videos presents challenges like dynamic sequences, static backgrounds, and illumination variations.
Purpose of the Study:
- To develop a novel deep learning model for accurate video gaze estimation.
- To address challenges in dynamic gaze tracking, background variations, and illumination changes.
- To incorporate personalization for improved gaze estimation with minimal data.
Main Methods:
- A deep learning model incorporating a spatial attention module to track spatial dynamics in videos.
- A temporal sequence model transforming spatial observations into temporal insights for accurate gaze prediction.
- Integration of Gaussian processes for individual-specific personalization using few labeled samples.
Main Results:
- Achieved state-of-the-art performance on the Gaze360 dataset, with a 2.5° improvement without personalization.
- Further improved performance by 0.8° with personalization using only three samples.
- Demonstrated efficacy in both within-dataset and cross-dataset evaluations.
Conclusions:
- The proposed spatial-temporal attention model significantly enhances video gaze estimation accuracy.
- Personalization with Gaussian processes allows for effective model adaptation with minimal data.
- The approach offers a robust solution for analyzing human attention and behavior from video data.

