Related Experiment Video
Updated: Dec 14, 2025

09:27
Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018
10.5K
Computational framework for fusing eye movements and spoken narratives for image annotation
Preethi Vaidyanathan1, Emily Prud'hommeaux2, Cecilia O Alm3
1Eyegaze Inc., Fairfax, VA, USA.
Journal of Vision
|July 18, 2020
Summary
This study integrates human gaze and spoken language to automatically label important image regions. This approach enhances computer vision by bridging the gap between machine processing and human understanding of visual data.
Area of Science:
- Computer Vision
- Human-Computer Interaction
- Natural Language Processing
- Cognitive Science
Background:
- Current computer vision systems struggle to replicate human-level image understanding.
- A gap exists between computational image processing and human perceptual interpretation.
- Integrating multimodal human data offers a pathway to bridge this understanding gap.
Purpose of the Study:
- To develop a framework that integrates human gaze and spoken language for image region labeling.
- To create meaningful mappings between eye movements and spoken descriptions for semantic annotation.
- To improve the accuracy of image region labeling by leveraging human perceptual data.
Main Methods:
- Utilized an unsupervised bitext alignment algorithm, originally for machine translation.
- Mapped participants' eye movements (gaze) to their spoken descriptions of images.
- Annotated image regions with linguistic labels based on these multimodal alignments.
Main Results:
- The proposed framework achieved higher accuracy in labeling image regions compared to baseline temporal alignments.
- Gaze-based clustering methods showed performance differences compared to image feature-based methods.
- Generated accurate multimodal alignments linking low-level image features with high-level semantic annotations.
Conclusions:
- Integrating human gaze and spoken language effectively labels perceptually important image regions.
- The framework provides a novel method for creating semantically rich image databases.
- The computational framework is adaptable to various multimodal data streams and visual domains.

