Related Experiment Video
Updated: Jan 5, 2026

07:36
Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
16.3K
Understanding Events by Eye and Ear: Agent and Verb Drive Non-anticipatory Eye Movements in Dynamic Scenes
Roberto G de Almeida1, Julia Di Nardo1, Caitlyn Antal1,2
1Department of Psychology, Concordia University, Montreal, QC, Canada.
Frontiers in Psychology
|October 26, 2019
Summary
This study on language and vision found that agent motion, not verb type, primarily guides attention in realistic scenes. Linguistic and visual processing interact later, using a common code.
Area of Science:
- Cognitive Science
- Psycholinguistics
- Visual Cognition
Background:
- Understanding the interplay between language and vision is crucial for explaining real-world human cognition.
- Previous research often used static or artificial stimuli, limiting insights into naturalistic interactions.
Purpose of the Study:
- To investigate how linguistic representations (verb semantics) and visual representations (agent motion) interact during the comprehension of dynamic, real-world scenes.
- To examine the timing and nature of this interaction using eye-tracking.
Main Methods:
- Participants viewed realistic video clips of everyday scenes while listening to synchronized descriptive sentences.
- Eye movements were monitored to track attention.
- Sentences varied in verb class (causative vs. perception) and scenes depicted agent motion (toward, away, neutral).
Main Results:
- Agent motion significantly influenced visual attention, with a stronger effect when the agent moved toward the target object.
- Verb semantics showed weaker effects, only modulating attention for causative verbs during motion toward the object.
- No anticipatory eye movements toward referents were observed based solely on verb type, unlike in studies with static scenes.
Conclusions:
- Early linguistic and visual processing in naturalistic settings may operate more independently than previously thought.
- Interaction occurs later, possibly at a central conceptual system utilizing a shared propositional code.
- Findings suggest a revised model for language-vision interaction in real-world contexts.

