Related Experiment Video
Updated: Sep 8, 2026

Investigating the Deployment of Visual Attention Before Accurate and Averaging Saccades via Eye Tracking and Assessment of Visual Sensitivity
Published on: March 18, 2019
Selective Visuo-Spatial Attention Revealed by Fine-grained Deep Visual Representations during Naturalistic Movie
Liting Wang1, Xin Zhang2, Pengcheng Xie3
1School of Automation, Northwestern Polytechnical University, Xi'an 710072, China; School of Information and Electrical Engineering, Hebei University of Engineering, Handan 056038, China.
Abstract:
Selective visuo-spatial attention is a complex and dynamic process that prioritizes relevant visual information while suppressing irrelevant distractors. Recent studies have demonstrated that incorporating computational visual attention models with functional magnetic resonance imaging (fMRI) in naturalistic paradigms (e.g., movie-watching) offers valuable insights into the neural mechanisms underlying selective visual attention. However, prior research has typically treated these models as "black boxes", overlooking the heterogeneity of their elementary hidden units, such as the convolutional filters in a convolutional neural network (CNN). In this study, we investigate whether decomposing a CNN-based computational visual attention model into its constituent filters can advance our understanding of the neural mechanisms of selective visual attention. Specifically, we treated each filter as a distinct saliency-related feature channel and derived its temporal responses during movie-clip processing. These temporal responses served as filter-specific visual attention regressors in general linear model (GLM) analysis applied to movie-watching fMRI data to localize associated brain regions. Our results demonstrate that these filters can be grouped into three clusters with distinct cortical activation patterns. Notably, the primary cluster reliably mapped onto canonical fronto-parietal attention networks alongside higher-order visual cortices, showing remarkable consistency with the global attention regressor and suggesting that decomposed filters successfully capture core visuospatial components. Crucially, the filter-specific analysis revealed distinct neural response patterns beyond the global saliency representation, involving auditory regions and the default mode network (DMN) during naturalistic viewing. Furthermore, eye-tracking analysis confirmed that the filter units effectively capture an auditory bias in visual attention. These findings offer valuable insights into the neural basis of fine-grained visual attention representations, particularly under naturalistic viewing conditions.

