Related Experiment Video
Updated: Oct 3, 2025

12:39
A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
Published on: January 18, 2020
7.8K
Convolutional Neural Network-Based Technique for Gaze Estimation on Mobile Devices.
Andronicus A Akinyelu1, Pieter Blignaut1
1Department of Computer Science and Informatics, Faculty of Natural and Agricultural Sciences, University of the Free State, Bloemfontein, South Africa.
Frontiers in Artificial Intelligence
|February 14, 2022
Summary
This study introduces a calibration-free Convolutional Neural Network (CNN) technique for accurate gaze estimation. Incorporating 39 facial landmarks significantly improves eye-tracking performance in real-world settings.
Area of Science:
- Computer Vision
- Human-Computer Interaction
- Machine Learning
Background:
- Current eye-tracking technologies are often costly and require explicit personal calibration, limiting their use in uncontrolled environments.
- Calibration procedures can be cumbersome, negatively impacting user experience and real-world applicability.
- There is a need for robust, calibration-free gaze estimation methods suitable for unconstrained settings.
Purpose of the Study:
- To develop and evaluate a novel Convolutional Neural Network (CNN) based technique for calibration-free gaze estimation.
- To improve the accuracy and usability of eye-tracking technology in real-world, unconstrained environments.
- To investigate the efficacy of incorporating facial landmark information into CNN models for gaze estimation.
Main Methods:
- A CNN-based approach was proposed, integrating a face component for feature extraction and a 39-point facial landmark component to encode eye shape and location.
- A comparative CNN model, accepting only face images, was developed for performance benchmarking.
- Experiments were conducted to compare the proposed technique against the baseline CNN model, including fine-tuning with the VGG16 pre-trained model.
Main Results:
- The proposed technique, utilizing both face and 39-point facial landmark components, demonstrated superior performance compared to the baseline CNN model.
- Fine-tuning the proposed technique with VGG16 further enhanced its performance, outperforming the fine-tuned baseline model.
- The inclusion of 39-point facial landmarks was shown to significantly improve the performance of CNN-based gaze estimation.
Conclusions:
- The proposed calibration-free CNN technique effectively enhances gaze estimation accuracy in unconstrained environments.
- The integration of 39-point facial landmarks is a viable strategy for improving the robustness and performance of eye-tracking systems.
- This approach offers a promising solution for more accessible and user-friendly eye-tracking applications.

