Related Experiment Video
Updated: May 26, 2026

07:09
Gaze in Action: Head-mounted Eye Tracking of Children's Dynamic Visual Attention During Naturalistic Behavior
Published on: November 14, 2018
Look Hear: Gaze Prediction for Speech-directed Human Attention.
Sounak Mondal1, Seoyoung Ahn2, Zhibo Yang3
1Stony Brook University, NY, USA.
Summary
This study introduces the Attention in Referral Transformer (ART) model to predict human attention during image-based object referral. ART accurately forecasts gaze patterns, improving human-computer interaction with spoken language.
Area of Science:
- Computer Vision
- Human-Computer Interaction
- Cognitive Science
Background:
- Effective human-computer interaction requires systems to understand how language influences user attention.
- Predicting user gaze is crucial for natural language understanding and interaction.
Purpose of the Study:
- To develop a model for incremental prediction of human attention during object referral tasks.
- To enhance natural language understanding by modeling word-level attention shifts.
Main Methods:
- Developed the Attention in Referral Transformer (ART) model, a multimodal transformer.
- Utilized an autoregressive transformer decoder for predicting fixation sequences.
- Created the RefCOCO-Gaze dataset with 19,738 human gaze scanpaths.
Main Results:
- ART outperforms existing methods in scanpath prediction accuracy.
- The model captures human attention patterns like waiting, scanning, and verification.
- Demonstrated the model's ability to jointly learn gaze behavior and grounding tasks.
Conclusions:
- ART provides a significant advancement in predicting human attention for referring expressions.
- The model's success suggests a pathway towards more intuitive human-computer dialogue systems.
- The RefCOCO-Gaze dataset enables further research in multimodal attention modeling.

