Related Experiment Video
Updated: Dec 23, 2025

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
Minimal videos: Trade-off between spatial and temporal information in human and machine vision
Guy Ben-Yosef1, Gabriel Kreiman2, Shimon Ullman3
1Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology, Cambridge, MA 02139, USA; Center for Brains, Minds and Machines, Massachusetts Institute of Technology, Cambridge, MA 02139, USA.
Human visual recognition efficiently integrates spatial and motion cues, even when each is insufficient alone. Current AI models struggle to replicate this full spatiotemporal interpretation, highlighting a gap in machine vision.
Area of Science:
- Cognitive Science
- Computer Vision
- Neuroscience
Background:
- Visual recognition typically relies on spatial or temporal information.
- Mechanisms for integrating space and time in vision are not well understood.
Purpose of the Study:
- To investigate how humans combine spatial and motion cues for visual recognition.
- To identify minimal video configurations for testing spatiotemporal integration.
- To compare human performance with state-of-the-art deep convolutional networks.
Main Methods:
- Analysis of "minimal videos" – short clips where recognition fails if spatial or temporal information is reduced.
- Human behavioral experiments using these minimal videos.
- Comparison of human recognition with deep convolutional network performance.
Main Results:
- Humans can recognize objects and actions by efficiently combining spatial and motion cues.
- Minimal videos require intact spatiotemporal information for recognition.
- Deep convolutional networks fail to replicate human recognition in these minimal video configurations.
Conclusions:
- Human vision effectively integrates spatial and temporal information for robust recognition.
- Current computational models lack critical mechanisms for full spatiotemporal interpretation.
- A significant gap exists between human and machine visual recognition capabilities.
Related Concept Videos
Depth Perception and Spatial Vision
Parallel Processing

