Related Experiment Video
Updated: Dec 29, 2025

04:43
Visualizing Visual Adaptation
Published on: April 24, 2017
9.5K
Scale and translation-invariance for novel objects in human vision.
Yena Han1, Gemma Roig2,3, Gad Geiger2
1Center for Brains, Minds and Machines, MIT, 77 Massachusetts Ave, Cambridge, MA, 02139, United States of America. yenahan@mit.edu.
Scientific Reports
|January 31, 2020
Summary
Humans demonstrate remarkable scale-invariance in object recognition after just one exposure. However, translation-invariance is limited, suggesting the visual system
Area of Science:
- Cognitive Neuroscience
- Computer Vision
- Human Visual Perception
Background:
- Object recognition is fundamental to human vision.
- Understanding invariance (tolerance to changes in scale, position, etc.) is crucial but challenging.
- Current computational models, like deep learning, often require extensive data for robust recognition.
Purpose of the Study:
- To investigate the extent of scale- and translation-invariance in one-shot learning of novel objects.
- To compare human visual system strategies with computational models, particularly deep learning architectures.
- To elucidate the neural computations underlying invariant object recognition.
Main Methods:
- Psychophysical experiments using Korean letters presented to naive subjects.
- Measuring recognition accuracy under varying scales and positions.
- Comparing experimental data with computational modeling of neural networks.
Main Results:
- Humans exhibit significant scale-invariance after a single exposure to a novel object.
- Translation-invariance is limited and depends on object size and position.
- Neural network models require explicit scale-invariance mechanisms (scale channels, eccentricity-dependent representations) to match human performance.
Conclusions:
- The human visual system achieves data-efficient, invariant object recognition through strategies distinct from current deep learning models.
- Incorporating scale-invariance and eccentricity-dependent representations is key for computational models.
- Eye movements play a critical role in the human visual system's invariant recognition capabilities.
Related Concept Videos
Depth Perception and Spatial Vision
1.7K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
1.7K
Perceptual Constancy
1.1K
Perceptual constancy is the ability to recognize that objects remain consistent and unchanged even when their appearance varies due to changes in sensory input. There are four main types of perceptual constancy: size constancy, shape constancy, color constancy, and brightness constancy.
Size constancy is the recognition that an object remains the same size, even when its image on the retina changes. For instance, a bus is perceived to be large enough to carry people, even if it looks tiny from...
Size constancy is the recognition that an object remains the same size, even when its image on the retina changes. For instance, a bus is perceived to be large enough to carry people, even if it looks tiny from...
1.1K

