雷达-摄像机融合在视角视图和鸟视图中用于3D物体检测
Yuhao Xiao1,2, Xiaoqing Chen1,2, Yingkai Wang1,2
1Chengdu Institute of Computer Application, Chinese Academy of Sciences, Chengdu 610213, China.
Sensors (Basel, Switzerland)
|October 16, 2025
概括
这项研究引入了一种用于3D物体检测的新型双视图融合方法,通过将雷达和相机视角视图融合,提高了深度估计的准确性. 这种方法与现有的鸟视融合方法相比,显著提高了检测性能.
科学领域:
- 计算机视觉 计算机视觉
- 传感器融合式传感器
- 机器人技术 机器人技术 机器人技术
背景情况:
- 使用雷达-摄像头融合的3D物体检测正在推进,因为它的成本效益,准确性和稳定性.
- 主导的鸟视图 (BEV) 融合范式依赖于精确的图像和雷达BEV特征.
- 从单眼图像中准确地估计深度,这对于图像BEV特征至关重要,仍然是一个具有挑战性的,错误的问题.
研究的目的:
- 通过融合相机和雷达视角视图 (PV) 功能来提高深度估计的准确性.
- 通过增强的深度估计来提高图像精度,BEV功能.
- 通过将精细的图像BEV特征与雷达BEV特征融合,实现更准确的3D对象检测.
主要方法:
- 开发了一种新的双视图融合范式,将摄像头和雷达视角视图结合起来.
- 设计了一个使用雷达截面 (RCS) 和深度信息进行准确的雷达数据投影的雷达图像生成模块.
- 实现了一种跨模式的功能融合模块,具有用于激光雷达和摄像头光伏功能的动态融合的注意力机制.
主要成果:
- 拟议的双视图融合范式在传统的BEV融合范式上表现出优越的性能.
- 在nuScenes 3D物体检测数据集上取得了最先进的结果.
- 实现了64.2 NDS和56.3 mAP,表明3D物体检测的显著改进.
结论:
- 双视图融合方法有效地提高了深度估计的准确性,从而改善了3D对象的检测.
- 在透视视图中融合雷达和摄像头数据,比仅仅依靠鸟视图融合提供了优势.
- 该方法为强大而准确的3D物体检测系统提供了有前途的进步.
相关概念视频
Depth Perception and Spatial Vision
1.8K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
1.8K
Relative Motion Analysis using Rotating Axes
876
Consider a component AB undergoing a linear motion. Along with a linear motion, point B also rotates around point A. To comprehend this complex movement, position vectors for both points A and B are established using a stationary reference frame.
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it...
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it...
876


