通过基于特征提取和集成的卷积囊网络检测RGB-D突出物体.
Kun Xu1,2,3, Jichang Guo4
1School of Electrical and Information Engineering, Tianjin University, Tianjin, 300000, People's Republic of China.
Scientific reports
|October 17, 2023
概括
这项研究引入了一个新的卷积囊网络用于突出物体检测,有效地解决了对象部分困境. 与现有算法相比,该新方法在降低计算需求的情况下实现了卓越的性能.
科学领域:
- 计算机视觉 计算机视觉
- 人工智能的人工智能
- 深度学习 (Deep Learning) 是一种深度学习.
背景情况:
- 完全卷积神经网络 (FCNN) 在使用RGB或RGB-D数据的突出对象检测方面表现出色,但在对象部分细分方面扎.
- 囊网络可以识别完整的对象,但是计算密集型和耗时的.
- 需要有效的方法,可以准确地分割突出的对象,同时保持对象部分关系.
研究的目的:
- 为突出物体检测提出一种新的卷积囊网络 (CNN),以解决对象部分困境.
- 与传统的囊网络相比,开发一种计算需求减少的方法.
- 为了提高突出物体细分的准确性和完整性.
主要方法:
- 使用VGG骨干用于RGB特征提取和集成.
- 整合了一个功能深度模块,将RGB功能与深度图像信息融合在一起.
- 采用功能集成的卷积囊网络,具有局部连接的路由,用于对象部分关系探索.
- 使用解卷性囊生成最终的突出地图.
主要成果:
- 提出的方法有效地解决了突出物体检测中的对象部分困境.
- 与23个最先进的算法相比,在四个RGB-D基准数据集上实现了卓越的性能.
- 证明了在保持高精度的同时减少了计算需求.
结论:
- 新型卷积囊网络为突出物体检测提供了有效的解决方案,克服了FCNN和传统囊网络的局限性.
- 该方法在细分完整性和准确性方面取得了显著的改进.
- 这种方法为突出的物体检测任务提供了一个计算效率高和高性能替代方案.
相关概念视频
Visual System
594
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
594
Vision
53.5K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
53.5K
Color Vision
593
Color perception begins in the retina, the light-sensitive layer at the back of the eye. Two main theories explain how colors are seen: the trichromatic theory and the opponent-process theory. The trichromatic theory, proposed by Thomas Young in 1802 and extended by Hermann von Helmholtz in 1852, suggests that color vision is based on three types of cone receptors in the retina. These cones are sensitive to different but overlapping ranges of wavelengths corresponding to red, blue, and green.
593
Force Classification
1.2K
Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
1.2K
Anatomy of the Eyeball
7.2K
The eye is a spherical, hollow structure composed of three tissue layers. The outer layer — the fibrous tunic, comprises the sclera — a white structure — and the cornea, which is transparent. The sclera encompasses some of the ocular surface, most of which is not visible. However, the 'white of the eye' is distinctively visible in humans compared to other species. The cornea, a clear covering at the front of the eye, enables light penetration. The eye's middle...
7.2K
Difference from Background: Limit of Detection
6.4K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
6.4K


