剩余视觉变压器和自适应融合自编码器用于单眼深度估计
Wei-Jong Yang1, Chih-Chen Wu2, Jar-Ferr Yang2
1Department of Artificial Intelligence and Computer Engineering, National Chin-Yi University of Technology, Taichung 411, Taiwan.
Sensors (Basel, Switzerland)
|January 11, 2025
概括
这项研究引入了一种用于单眼深度估计的新型自动编码器,通过单个摄像头增强3D传感. 该模型在深度地图预测中实现了卓越的准确性和减少错误,改善了自动驾驶等应用程序.
科学领域:
- 计算机视觉 计算机视觉
- 深度学习 (Deep Learning) 是一种深度学习.
- 3D 感应 3D 感应
背景情况:
- 单眼深度估计对于3D场景重建和自动驾驶等应用至关重要.
- 深度学习的进步使单眼深度估计能够超越传统的立体相机系统.
- 从单视图图像中准确的深度感知仍然是一个重大挑战.
研究的目的:
- 提出一个端到端监督的单眼深度估计自编码器,使用单个摄像头.
- 通过一种新的网络架构来提高深度图的精度.
- 为了提高前景对象的深度估计性能.
主要方法:
- 开发了一种带有混合卷积神经网络和视觉转换器编码器的自动编码器.
- 实现了自适应融合解码器,以实现有效的功能合并.
- 在训练中使用了与人类感知对齐的损失函数.
主要成果:
- 拟议的自动编码器有效地从单视图彩色图像预测深度图.
- 与现有方法相比,第一个准确率提高了28%.
- 在纽约大学数据集上,将根平均平方误差降低了约27%.
结论:
- 开发的自动编码器展示了高精度单眼深度估计能力.
- 混合编码器和自适应解码器架构有效地捕获多尺度特征.
- 感知对齐的损失函数改善了对关键前景对象深度的关注.
相关概念视频
Depth Perception and Spatial Vision
548
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
548
Vision
52.9K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
52.9K


