Related Experiment Video
Updated: Jun 7, 2025

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
471
Teaching deep networks to see shape: Lessons from a simplified visual world
Christian Jarvers1, Heiko Neumann1
1Institute for Neural Information Processing, Ulm University, Ulm, Germany.
Plos Computational Biology
|November 11, 2024
Summary
Deep neural networks struggle to model primate vision because they over-rely on color and texture over shape. New research reveals that specific learning algorithms can improve deep networks
Area of Science:
- Computational neuroscience
- Artificial intelligence
- Computer vision
Background:
- Deep neural networks (DNNs) are successful models of the primate visual system.
- DNNs exhibit a strong shape-dependence in human vision, unlike current models.
- Humans prioritize shape for category judgments, while DNNs favor color and texture.
Purpose of the Study:
- Investigate why DNNs fail to capture the shape-dependence of primate vision.
- Identify the underlying reasons for DNNs' bias towards non-shape features.
- Propose solutions to enhance DNNs' sensitivity to shape.
Main Methods:
- Designed artificial image datasets with isolated shape, color, and texture features.
- Trained DNNs from scratch on these datasets using single features and combinations.
- Analyzed network architectures and learning algorithms, specifically mini-batch gradient descent.
Main Results:
- Some DNN architectures were unable to learn shape features effectively.
- Other architectures showed a bias towards color and texture, despite being capable of learning shape.
- This bias was linked to weight update interactions during mini-batch gradient descent.
Conclusions:
- Current DNNs and learning algorithms are not optimized for shape-based visual processing.
- Mini-batch gradient descent contributes to the bias against shape features.
- Developing learning algorithms with sparser, more local weight changes is crucial for improving DNNs' shape sensitivity and modeling human vision.
Related Concept Videos
Depth Perception and Spatial Vision
593
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
593
Gestalt Principles of Perception
279
Gestalt principles provide a framework for understanding how humans perceive objects as unified wholes within their context. These principles are essential in explaining the cognitive processes that make sense of complex visual stimuli by organizing them into coherent groups. One fundamental principle is proximity, which posits that objects located close to each other are perceived as a collective group. For instance, when dots are positioned near one another, the visual system interprets them...
279
Visual System
551
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
551
Visual Agnosia
176
Visual agnosia is a condition characterized by the inability to recognize visually presented objects despite having normal vision. For instance, a person with visual agnosia can describe the shape and color of an object but cannot identify or name it. This impairment does not affect their visual field, acuity, color vision, brightness discrimination, language, or memory. An example of this condition in a social setting is someone at a dinner party asking for "that silver thing with a round...
176
Vision
52.9K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
52.9K
Role of Shaping in Operant Conditioning
267
Shaping is a technique used in operant conditioning to train complex behaviors by rewarding successive approximations toward the target behavior. This method is necessary because organisms are unlikely to perform complex behaviors spontaneously. Instead, shaping breaks down the desired behavior into small, manageable steps.
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
The steps involved in shaping begin with reinforcing any response that resembles the desired behavior. For example, parents might praise a child for picking up one toy. As...
267

