Related Experiment Video
Updated: Jun 19, 2025

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
515
C2BG-Net: Cross-modality and cross-scale balance network with global semantics for multi-modal 3D object detection
Bonan Ding1, Jin Xie1, Jing Nie2
1School of Big Data and Software Engineering, Chongqing University, Chongqing, 400044, China.
Summary
This study introduces novel networks for multi-modal 3D object detection, improving feature aggregation by balancing sensor data significance. The proposed methods enhance accuracy in autonomous driving applications.
Area of Science:
- Computer Vision
- Robotics
- Artificial Intelligence
Background:
- Multi-modal 3D object detection fuses RGB images and lidar point-clouds for autonomous driving.
- Existing methods use simple feature aggregation (addition/multiplication), struggling to balance modal significance and potentially introducing noise.
- Image features at different levels exhibit receptive field imbalances, hindering comprehensive analysis.
Purpose of the Study:
- To develop advanced networks for effective multi-modal feature aggregation in 3D object detection.
- To address the challenges of balancing significance between RGB and lidar data.
- To overcome receptive field imbalances in multi-level image features for improved detection.
Main Methods:
- Proposed two novel networks: Cross-Modality Network (CMN) and Cross-Scale Network (CSN).
- CMN utilizes cross-modality attention and an auxiliary 2D detection head for balanced feature significance.
- CSN employs cross-scale attention to reconcile differing receptive fields across image levels. Introduced Local with Global Voxel Attention Encoder (LGVAE) for enhanced point-level feature extraction.
Main Results:
- The proposed CMN and CSN networks demonstrated consistent improvements across various 3D object detection frameworks.
- Achieved a significant 3.1% absolute gain over the baseline MVXNet on the challenging Dense Hard set.
- Validated effectiveness and versatility on KITTI, Dense, and nuScenes benchmarks.
Conclusions:
- The developed CMN and CSN networks offer a superior approach to multi-modal feature aggregation.
- The methods effectively balance modal significance and address receptive field disparities, enhancing 3D object detection performance.
- The proposed techniques show significant promise for advancing autonomous driving perception systems.

