Related Experiment Video
Updated: May 2, 2026

13:51
Cross-Modal Multivariate Pattern Analysis
Published on: November 9, 2011
21.0K
Investigating the temporal dynamics and modeling of mid-level feature representations in humans
Agnessa Karapetian1,2,3, Alexander Lenders1, Vanshika Bawa4
1Department of Education and Psychology, Freie Universität Berlin, Berlin, Germany.
Imaging Neuroscience (Cambridge, Mass.)
|May 1, 2026
Summary
Mid-level visual features, crucial for perception, are processed between low- and high-level features (~100-250 ms). This study introduces 3D-rendered stimuli to investigate these intermediate visual processing stages in humans and artificial neural networks.
Area of Science:
- Neuroscience
- Computer Vision
- Cognitive Science
Background:
- Visual perception involves a hierarchy from low-level (edges) to high-level (object categories) features.
- The processing of intermediate, or mid-level, visual features remains poorly understood.
- Understanding mid-level vision is key to bridging sensory input and semantic understanding.
Purpose of the Study:
- To investigate the timing and nature of mid-level visual feature processing in the human brain.
- To introduce and validate a novel stimulus set of 3D-rendered images and videos for studying mid-level vision.
- To compare human mid-level visual processing with convolutional neural network (CNN) models.
Main Methods:
- Developed a stimulus set of 3D-rendered naturalistic images and videos with annotations for low-, mid-, and high-level features.
- Recorded electroencephalography (EEG) responses during stimulus presentation.
- Trained linearized encoding models to predict EEG responses from stimulus annotations.
Main Results:
- Mid-level features (reflectance, depth, normals, lighting, skeleton) were processed between approximately 100 and 250 ms post-stimulus.
- This timing places mid-level feature processing between low- and high-level features, suggesting a bridging role.
- CNNs showed similar processing order for mid-level features in videos, but not for low- or high-level features, compared to humans.
Conclusions:
- Mid-level visual features are critical for surface and shape processing, linking sensory and semantic stages.
- The 3D-rendered stimulus set is a valuable tool for future research on mid-level vision in both biological and artificial systems.
- The findings provide insights into the hierarchical organization of the human visual system and the capabilities of CNNs.
