Related Experiment Video
Updated: Sep 2, 2026

End-To-End Deep Neural Network for Salient Object Detection in Complex Environments
Published on: December 15, 2023
Shared texture-like representations underlie deep neural network alignment with human visual processing
Jessica Loke1, Lynn K A Sörensen2, Iris I A Groen3
1Department of Psychology, University of Amsterdam, Nieuwe Achtergracht 129-B, 1018WS Amsterdam, the Netherlands; Amsterdam Brain & Cognition (ABC) Center, University of Amsterdam, Nieuwe Achtergracht 129-B, 1018WS Amsterdam, the Netherlands.
Abstract:
Deep neural networks (DNNs) excel at predicting neural responses across the visual hierarchy,1,2,3,4,5 a success widely interpreted as evidence of shared object recognition computations.6,7 Yet improving DNN object recognition accuracy does not reliably increase neural predictivity,8,9 and even untrained networks predict brain responses above chance.9,10,11 This disconnect suggests that object recognition may not drive DNN-brain alignment. Texture-like statistics are represented in both DNNs and mid-level visual cortex (A.V. Jagadeesh and M. Livingstone, 2024, ICLR, presentation).12,13,1416 In natural images, these statistics are carried by objects and backgrounds, shaping representations and recognition in both systems.16,17,18,19 Does DNN-brain alignment reflect a shared sensitivity to object-related information or texture-like statistics? To dissociate these factors, we recorded electroencephalograms (EEGs) from 57 participants viewing natural scenes, texture-synthesized images preserving local statistics while disrupting global form, and object-only images with backgrounds removed. If alignment reflects texture-like statistics, then it should peak for texture-synthesized images. If it reflects object-related processing, then alignment should be strongest for natural and object-only conditions, which preserve object information. We compared EEG responses with DNN activations via weighted representational similarity analysis.20,21 Texture-synthesized images yielded the strongest DNN-EEG alignment, peaking in early responses (<200 ms) and explaining up to ∼85% of noise-ceiling-normalized explainable variance versus ∼44% for natural and ∼55% for isolated objects. Crucially, object categories were more decodable for natural and object-only images than texture-synthesized images, yet these object-rich conditions showed weaker alignment. This dissociation reveals that DNNs capture the texture-statistical component of early visual responses while failing to explain later, object-related variance.
Related Concept Videos
Parallel Processing
Vision
Depth Perception and Spatial Vision
