视觉数值感应真实世界的场景,由深度神经网络和人类共享
Wu Wencheng1,2, Yingxi Ge3,4,5, Zhentao Zuo2,3,4,5
1AHU-IAI AI Joint Laboratory, Anhui University, Hefei, 230601, China.
Heliyon
|August 10, 2023
概括
深度神经网络 (DNN) 可以在现实世界的场景中感知人数. DNN中的一个群组编码机制,特别是AlexNet,揭示了场景特征如何影响人类的数字感知.
科学领域:
- 认知科学 认知科学
- 计算机视觉 计算机视觉
- 神经科学是一个神经科学.
背景情况:
- 深度神经网络 (DNN) 已经显示出视觉数感能力.
- 以前的研究经常使用简单的几何图形,让现实世界场景的容量不清楚.
研究的目的:
- 在现实场景中使用DNN调查数字感知.
- 探索场景特征如何影响人类的数值感知.
主要方法:
- 利用AlexNet分析场景图像中的数字感知.
- 在类别层单位中检查了组激活模式.
- 开发了一种解码方法来提取数量信息.
主要成果:
- 多样性是由DNN中的组激活模式表示的.
- 全球激活随着对象数量的增加而增加;激活变量减少.
- 场景嵌入系数预测对象对数值感知的贡献.
- 在DNN和人类中观察到高嵌入系数图像的优化性能.
结论:
- 通过群组激活模式,DNN代表了场景中的多样性.
- 通过DNN识别的场景特征可以调节人类的感知.
- 一个群组编码机制是这种视觉数感现象的基础.
相关概念视频
Depth Perception and Spatial Vision
720
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
720
Visual System
617
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
617
Vision
53.6K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
53.6K
Parallel Processing
181
The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
181
Perceptual Constancy
441
Perceptual constancy is the ability to recognize that objects remain consistent and unchanged even when their appearance varies due to changes in sensory input. There are four main types of perceptual constancy: size constancy, shape constancy, color constancy, and brightness constancy.
Size constancy is the recognition that an object remains the same size, even when its image on the retina changes. For instance, a bus is perceived to be large enough to carry people, even if it looks tiny from...
Size constancy is the recognition that an object remains the same size, even when its image on the retina changes. For instance, a bus is perceived to be large enough to carry people, even if it looks tiny from...
441
Visual Agnosia
237
Visual agnosia is a condition characterized by the inability to recognize visually presented objects despite having normal vision. For instance, a person with visual agnosia can describe the shape and color of an object but cannot identify or name it. This impairment does not affect their visual field, acuity, color vision, brightness discrimination, language, or memory. An example of this condition in a social setting is someone at a dinner party asking for "that silver thing with a round...
237


