Related Experiment Video
Updated: Oct 10, 2026

End-To-End Deep Neural Network for Salient Object Detection in Complex Environments
Published on: December 15, 2023
A Depth Complementary Framework for Monocular 3D Object Detection in Autonomous Driving
Abstract:
Monocular 3D object detection offers a cost-effective solution for 3D localization from a single image in autonomous driving, yet its accuracy remains fundamentally limited by depth estimation due to the ill-posed nature of 2D-to-3D mapping. Existing methods attempt to reduce estimation variance by ensembling predictions from multiple depth branches that primarily rely on local cues. However, we observe that these estimates exhibit systematic one-sided biases, where depth prediction errors for the same object tend to be predominantly either positive or negative. To address this limitation, we propose MonoCD, a novel framework that enhances depth complementarity through two key innovations. First, we introduce a depth branch that incorporates ground elevation as a global cue to complement conventional local cues. Second, we develop a bias-compensating formulation for depth prediction that applies an opposite-signed design to local cues, actively counteracting their systematic biases. Owing to these two components, MonoCD achieves substantially improved depth complementarity by generating more anti-correlated estimates. Extensive experiments on the KITTI and Waymo benchmarks show that MonoCD achieves state-of-the-art performance. Moreover, the proposed depth branch is lightweight and plug-and-play, making it readily applicable to boost existing monocular 3D detection frameworks.