几何Former:半卷积变压器与几何感知集成,用于自动驾驶场景中的深度完成
1National Key Laboratory of Automotive Chassis Integration and Bionics, Jilin University, Changchun 130025, China.
Sensors (Basel, Switzerland)
|January 8, 2025
概括
这项研究引入了一种用于自动驾驶的深度完成的新方法,通过融合视觉变压器和卷积来提高准确性. 新方法显著改善了边缘和透明区域的恢复,实现了最先进的结果.
科学领域:
- 计算机视觉 计算机视觉
- 机器人技术 机器人技术 机器人技术
- 自动驾驶自动驾驶的自动驾驶
背景情况:
- 深度完成对于自动驾驶中的同时定位和映射 (SLAM) 和运动结构 (SfM) 是至关重要的.
- 视觉变压器 (ViT) 和卷积融合方法已经提高了深度完成精度.
- 现有的方法在复杂场景中的细节恢复和不完整的融合方面扎.
研究的目的:
- 通过解决细节恢复和特征融合方面的局限性,提高深度完成精度.
- 在深度地图中改善对几何结构,边缘和透明区域的感知.
- 为了实现最先进的性能,深入完成任务.
主要方法:
- 提出了一种半卷积视觉变压器,以优化局部连续性.
- 设计了一个几何感知模块,用于学习空间相关性和几何特征.
- 引入了一种新的双阶段融合策略,具有可学习的信心,用于改进功能集成.
主要成果:
- 在NYU-Depth-v2和KITTI深度完成数据集上实现了最先进的 (SoTA) 性能.
- 在NYU-Depth-v2数据集上实现了创纪录的低根平均平方误差 (RMSE) 87.9毫米.
- 在现实世界的道路场景中展示了概括能力.
结论:
- 拟议的方法有效地通过优化局部连续性和几何感知来增强深度完成.
- 双阶段的融合策略显著改善了特征集成,减少了异常值和波纹.
- 该模型代表了自动驾驶应用的深度完成的重大进步.
相关概念视频
Depth Perception and Spatial Vision
548
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
548
Transformers with Off-Nominal Turns Ratios
138
In scenarios involving parallel transformers with disparate ratings, developing per-unit models requires accommodating off-nominal turns ratios. This situation arises when the selected base voltages are not proportional to the transformer’s voltage ratings. Consider a transformer where the rated voltages are related by the term a. If the chosen voltage bases satisfy a relationship involving term b, term c is defined as the ratio of these bases. This ratio is then substituted into the...
138
Parallel Processing
145
The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
145
Visual System
501
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
501
Transformers in Distribution System
98
Transformers in distribution systems can be broadly categorized into distribution substation transformers and other distribution transformers. They are crucial for stepping down high transmission voltages to levels suitable for distribution and end-user applications.
Distribution substation transformers come in various ratings and typically use mineral oil for insulation and cooling. To prevent moisture and air from entering the oil, some transformers use an inert gas like nitrogen to fill the...
Distribution substation transformers come in various ratings and typically use mineral oil for insulation and cooling. To prevent moisture and air from entering the oil, some transformers use an inert gas like nitrogen to fill the...
98
Convolution: Math, Graphics, and Discrete Signals
226
In any LTI (Linear Time-Invariant) system, the convolution of two signals is denoted using a convolution operator, assuming all initial conditions are zero. The convolution integral can be divided into two parts: the zero-input or natural response and the zero-state or forced response, with t0 indicating the initial time.
To simplify the convolution integral, it is assumed that both the input signal and impulse response are zero for negative time values. The graphical convolution process...
To simplify the convolution integral, it is assumed that both the input signal and impulse response are zero for negative time values. The graphical convolution process...
226


