Look Hear: Gaze Prediction for Speech-directed Human Attention.

Sounak Mondal1, Seoyoung Ahn2, Zhibo Yang3

  • 1Stony Brook University, NY, USA.

Computer Vision - ECCV ... : ... European Conference on Computer Vision : Proceedings. European Conference on Computer Vision
|May 25, 2026
PubMed
Summary

This study introduces the Attention in Referral Transformer (ART) model to predict human attention during image-based object referral. ART accurately forecasts gaze patterns, improving human-computer interaction with spoken language.