在深度视觉模型中对不变的实证评估
Konstantinos Keremis1, Eleni Vrochidou1, George A Papakostas1
1MLV Research Group, Department of Informatics, Democritus University of Thrace, 65404 Kavala, Greece.
Journal of imaging
|September 26, 2025
概括
深度学习模型与图像旋转和尺度变化作斗争. 视觉变压器 (ViT) 在识别任务中表现出比卷积神经网络 (CNN) 更好的模糊和噪声强度.
科学领域:
- 计算机视觉 计算机视觉
- 深度学习 (Deep Learning) 是一种深度学习.
- 机器学习 机器学习
背景情况:
- 深度学习模型需要强大,以对现实世界的应用程序进行图像变化.
- 不变量,就像处理转换一样,对于可靠的计算机视觉系统至关重要.
研究的目的:
- 实证地评估现代卷积神经网络 (CNN) 和视觉转换器 (ViT) 的图像不变性.
- 在对象本地化,识别和语义细分任务中评估模型的稳定性.
主要方法:
- 在基准数据集 (COCO,ImageNet) 上测试了30个CNN和ViT模型.
- 引入受控扰动 (模糊,噪音,旋转,尺度) 来评估强度.
- 使用mIoU和分类准确性 (Acc) 等指标来量化性能下降.
主要成果:
- 对于识别任务,ViT在模糊和噪音方面表现优于CNN.
- 无论是CNN还是ViT都显示出对旋转和极端规模转换的脆弱性.
- 语义细分模型,特别是SegFormer和Mask2Former,对几何变化的弹性更高.
结论:
- 当前的深度学习模型对常见的图像转换具有重大漏洞.
- 分段模型显示出强度的承诺,但需要进一步研究旋转和规模不变性.
- 发现挑战了假设,并为开发更强大的视觉系统提供了洞察力.
相关概念视频
Perceptual Constancy
1.3K
Perceptual constancy is the ability to recognize that objects remain consistent and unchanged even when their appearance varies due to changes in sensory input. There are four main types of perceptual constancy: size constancy, shape constancy, color constancy, and brightness constancy.
Size constancy is the recognition that an object remains the same size, even when its image on the retina changes. For instance, a bus is perceived to be large enough to carry people, even if it looks tiny from...
Size constancy is the recognition that an object remains the same size, even when its image on the retina changes. For instance, a bus is perceived to be large enough to carry people, even if it looks tiny from...
1.3K
Depth Perception and Spatial Vision
1.8K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
1.8K
Difference from Background: Limit of Detection
8.0K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
8.0K


