Related Experiment Video
Updated: Jun 24, 2026

03:31
End-To-End Deep Neural Network for Salient Object Detection in Complex Environments
Published on: December 15, 2023
Ecological Vision Hypothesis: Training Deep Neural Networks for Robustness and Human Alignment.
Frank Tong1,2, Hojin Jang3
11Department of Psychology, Vanderbilt University, Nashville, Tennessee, USA;
Annual Review of Vision Science
|June 22, 2026
Summary
Deep neural networks (DNNs) show promise but lack human vision robustness. Training DNNs with ecologically relevant challenges, like blur, could improve their flexibility and alignment with human visual perception.
Area of Science:
- Neuroscience
- Computer Vision
- Cognitive Science
Background:
- Deep neural networks (DNNs) are advanced neurocomputational models of the visual system.
- While DNNs excel at predicting neural responses to clear images, they exhibit brittleness under ambiguous viewing conditions.
- Human vision demonstrates remarkable robustness to challenges like noise, blur, and occlusion, unlike typical DNNs.
Purpose of the Study:
- To investigate the ecological vision hypothesis for improving DNN robustness.
- To explore how training DNNs with ecologically relevant visual challenges enhances human-like visual processing.
- To understand the role of blur in visual perception and its implications for DNN development.
Main Methods:
- Discussing the ecological vision hypothesis.
- Analyzing the limitations of current DNNs in handling visual ambiguities.
- Proposing training strategies for DNNs based on human visual experience.
Main Results:
- DNNs trained on standard datasets lack the robustness of human vision.
- Challenging viewing conditions, prevalent in natural environments, are key to acquiring robust vision.
- Blur may shift visual processing from local texture to global shape sensitivity.
Conclusions:
- The ecological vision hypothesis suggests training DNNs with naturalistic challenges can improve robustness.
- DNNs trained with ecologically relevant data, particularly 3D scene and shape information, are expected to align better with human vision.
- Addressing DNN brittleness requires incorporating the complexities of real-world visual input.