Related Experiment Videos
Exploiting audio-visual modalities in videos: Object detection via multi-stage bilateral coupling network.
Qifeng Liu1, Zujun Yu2, Liqiang Zhu2
1State Key Laboratory of Advanced Rail Autonomous Operation, Beijing Jiaotong University, Beijing, 100044, Beijing, China; School of Mechanical, Electronic and Control Engineering, Beijing Jiaotong University, Beijing, 100044, Beijing, China.
Summary
The Multi-Stage Bilateral Coupling Network (MSBCNet) advances audio-visual object detection by treating audio and visual data equally during training and inference. This novel approach enables the detection of both sounding and non-sounding objects, improving comprehensive perception.
Area of Science:
- Computer Vision
- Machine Learning
- Signal Processing
Background:
- Existing audio-visual object detection methods often use asymmetric fusion, neglecting one modality during inference.
- Current methods primarily focus on sound source localization, limiting detection to sounding objects.
- This prevents the full exploitation of integrated audio-visual data for comprehensive perception.
Purpose of the Study:
- To develop an object detection method that effectively utilizes both audio and visual modalities as co-equal partners.
- To enable the detection of both sounding and non-sounding objects in video data.
- To introduce the Multi-Stage Bilateral Coupling Network (MSBCNet) for holistic audio-visual learning.
Main Methods:
- Proposed the Multi-Stage Bilateral Coupling Network (MSBCNet) with a multi-stage guided mechanism (coarse perception, modality fusion, fine localization).
- Employed a video teacher and audio-visual student network design, utilizing feature consistency losses and a Cross-modal Feature Aggregation Module (CMFAM).
- Incorporated contrastive distribution alignment loss and a Salient Audio-Visual Response Module (SAVRM) to strengthen audio-visual coupling.
Main Results:
- MSBCNet achieved 87.38% mAP on a public multimodal dataset and 81.42% mAP on a railway dataset.
- Demonstrated significant superiority in audio-visual cross-modal object detection compared to existing methods.
- Ablation studies confirmed the efficacy of individual components within the MSBCNet framework.
Conclusions:
- MSBCNet is the first framework to holistically leverage audio and visual modalities as co-equal partners for object detection during both training and inference.
- The proposed method effectively addresses the limitations of existing approaches, enabling detection of both sounding and non-sounding objects.
- MSBCNet significantly advances the field of audio-visual cross-modal object detection.