Related Experiment Video
Updated: Jun 6, 2026

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
9.0K
Quantifying Dwell Time With Location-based Augmented Reality: Dynamic AOI Analysis on Mobile Eye Tracking Data With
Julien Mercier1,2, Olivier Ertz1, Erwan Bocher2
1MEI, School of Engineering and Management Vaud, HES-SO, Switzerland.
Journal of Eye Movement Research
|June 12, 2024
Summary
Researchers developed a Vision Transformer (ViT) model to accurately analyze mobile eye-tracking data from naturalistic studies. This automated method overcomes limitations of manual annotation for moving targets, improving data analysis efficiency.
Area of Science:
- Computer Vision
- Human-Computer Interaction
- Cognitive Science
Background:
- Mobile eye tracking offers valuable egocentric vision data for naturalistic studies but suffers from noise and analysis challenges with moving targets.
- Current analysis tools are limited, often necessitating time-consuming manual annotation, hindering the scalability of mobile eye-tracking research.
- Nonlinear movement and object disappearances in outdoor settings complicate automated area of interest analysis.
Purpose of the Study:
- To introduce an automated method for analyzing mobile eye-tracking data, specifically addressing challenges with moving targets in naturalistic environments.
- To improve the efficiency and accuracy of processing noisy eye-tracking data from outdoor studies.
- To enable more scalable and less labor-intensive analysis of mobile eye-tracking data.
Main Methods:
- A fine-tuned Vision Transformer (ViT) model was developed for classifying frames containing gaze markers.
- The ViT model was trained on a manually labeled subset (1.98%) of the entire dataset over three epochs.
- The model's performance was evaluated using hold-out data to assess its accuracy.
Main Results:
- The fine-tuned Vision Transformer model achieved a high accuracy of 99.34% on hold-out data.
- The method was successfully applied to quantify participant dwell time on a tablet during an outdoor augmented reality biodiversity education application test.
- Demonstrated the model's capability to handle complex, naturalistic scenarios with moving elements.
Conclusions:
- The developed Vision Transformer-based method offers an accurate and efficient solution for analyzing mobile eye-tracking data in naturalistic studies.
- This approach significantly reduces reliance on manual annotation, paving the way for broader application of mobile eye tracking.
- The method shows potential for application in diverse research areas requiring analysis of gaze behavior in dynamic environments.

