Related Experiment Video
Updated: Mar 15, 2026

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
Deep Networks Can Resemble Human Feed-forward Vision in Invariant Object Recognition
Saeed Reza Kheradpisheh1,2, Masoud Ghodrati3,4, Mohammad Ganjtabesh1
1Department of Computer Science, School of Mathematics, Statistics, and Computer Science, University of Tehran, Tehran, Iran.
Deep convolutional neural networks (DCNNs) show promise in object recognition. Deeper networks match human performance and representations better with increased viewpoint variations, with very deep nets even surpassing human capabilities.
Area of Science:
- Computer Vision
- Cognitive Science
- Neuroscience
Background:
- Deep convolutional neural networks (DCNNs) mimic human visual processing with hierarchical feature extraction.
- Previous studies have not fully explored DCNN performance against human object recognition, especially concerning viewpoint variations.
- The similarity in architecture between DCNNs and the human visual system warrants investigation into functional parallels.
Purpose of the Study:
- To compare the performance of state-of-the-art DCNNs, HMAX, and shallow models against human object recognition under varying viewpoint conditions.
- To determine if DCNNs exhibit human-like error patterns and representations.
- To assess the impact of viewpoint variation magnitude on model and human performance.
Main Methods:
- Benchmarking eight DCNNs, the HMAX model, and a shallow model against human performance using backward masking.
- Systematically controlling the magnitude of viewpoint variations in object recognition tasks.
- Analyzing performance metrics, error distributions, and representational consistency with human behavior.
Main Results:
- Shallow networks outperformed deep networks and humans at low viewpoint variations.
- Deeper networks were required to match human performance and error distributions with increasing viewpoint variations.
- An 18-layer DCNN surpassed human performance and demonstrated the most human-like representations at high variation levels.
Conclusions:
- The depth of neural networks is crucial for achieving human-level view-invariant object recognition, particularly with significant viewpoint changes.
- DCNNs can replicate human-like representations and error patterns, especially deeper architectures.
- Network depth is a key factor in bridging the gap between artificial and human visual recognition capabilities.
Related Concept Videos
Vision
Visual System
Once through the pupil, the light passes through the lens, a...
Association Areas of the Cortex
Prefrontal Association Area: This area is located in the frontal lobe and is involved in planning, decision-making, and moderating social behavior. It connects with primary motor areas,...
Parallel Processing
Depth Perception and Spatial Vision
Color Vision
