通过在卷积神经网络中结合强度以模糊,改进了人类视觉的建模
1Department of Psychology and Vanderbilt Vision Research Center Vanderbilt University.
bioRxiv : the preprint server for biology
|August 14, 2023
概括
用模糊图像训练卷积神经网络 (CNN) 提高了它们识别物体的能力,使它们对视觉噪声更具稳定性,并更好地与人类感知保持一致.
科学领域:
- 神经科学是一个神经科学.
- 计算机视觉 计算机视觉
- 人工智能的人工智能
背景情况:
- 人类视觉系统有效地处理退化的视觉输入,这是人工神经网络 (ANN) 培训中经常被忽视的因素.
- 标准卷积神经网络 (CNN) 可能过度依赖高频细节,导致与生物视觉的性能差异.
研究的目的:
- 研究模糊训练数据对CNN表现及其与生物视觉对应性的影响.
- 测试假设,没有模糊训练的CNN系统地偏离了人类的感知.
- 通过结合模糊图像训练来增强CNN的稳定性和形状敏感性.
主要方法:
- 标准CNN与在清晰和模糊图像上训练的CNN的比较.
- 评估CNN在各种观看条件下预测神经反应的能力.
- 评估CNN对塑造信息的敏感性和对视觉噪声的强度.
主要成果:
- 用模糊图像训练的CNN在预测神经反应方面超过了标准的CNN.
- 经过模糊训练的CNN表现出对塑造信息的敏感性增加.
- 这些CNN对各种形式的视觉噪声表现出更强的稳定性,更好地与人类的感知保持一致.
结论:
- 将模糊的视觉数据纳入培训对于开发更具生物可信性和强大的CNN至关重要.
- 这种神经计算方法强调了视觉退化在塑造有效的物体识别系统中的重要性.
相关概念视频
Convolution Properties II
233
The important convolution properties include width, area, differentiation, and integration properties.
The width property indicates that if the durations of input signals are T1 and T2, then the width of the output response equals the sum of both durations, irrespective of the shapes of the two functions. For instance, convolving two rectangular pulses with durations of 2 seconds and 1 second results in a function with a width of 3 seconds.
The area property asserts that the area under the...
The width property indicates that if the durations of input signals are T1 and T2, then the width of the output response equals the sum of both durations, irrespective of the shapes of the two functions. For instance, convolving two rectangular pulses with durations of 2 seconds and 1 second results in a function with a width of 3 seconds.
The area property asserts that the area under the...
233
Visual System
617
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
617
Focusing of Light in the Eye
2.9K
Light rays enter the eye through the cornea, a transparent dome-shaped tissue that is the eye's outermost layer. The cornea bends or refracts, light rays traveling to the pupil. The shape of the cornea determines how much of the light is bent and whether the image will be focused correctly on the retina at the back of the eye. Once the light has passed through both refraction layers, it converges into a single focal point onto a small area. This is where photoreceptors start transforming...
2.9K
Deconvolution
188
Deconvolution, also known as inverse filtering, is the process of extracting the impulse response from known input and output signals. This technique is vital in scenarios where the system's characteristics are unknown, and they must be inferred from the observable signals.
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
188
Convolution Properties I
180
Convolution computations can be simplified by utilizing their inherent properties.
The commutative property reveals that the input and the impulse response of an LTI (Linear Time-Invariant) system can be interchanged without affecting the output:
The commutative property reveals that the input and the impulse response of an LTI (Linear Time-Invariant) system can be interchanged without affecting the output:
180
Depth Perception and Spatial Vision
720
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
720


