Related Experiment Videos
Exploiting audio-visual modalities in videos: Object detection via multi-stage bilateral coupling network
Qifeng Liu1, Zujun Yu2, Liqiang Zhu2
1State Key Laboratory of Advanced Rail Autonomous Operation, Beijing Jiaotong University, Beijing, 100044, Beijing, China; School of Mechanical, Electronic and Control Engineering, Beijing Jiaotong University, Beijing, 100044, Beijing, China.
Abstract:
Despite the fact that video data integrates visual and auditory modalities to support comprehensive perception, existing audio-visual object detection methods predominantly employ asymmetric fusion paradigms, using one modality as an auxiliary cue during training while neglecting it during inference. Methods leveraging both modalities at inference typically focus on sound source localization, detecting only sounding objects. This limitation prevents the full exploitation of video data. To construct an object detection method applicable to both sounding and non-sounding objects, we innovatively propose the Multi-Stage Bilateral Coupling Network (MSBCNet) - the first framework to holistically leverage audio and visual modalities as co-equal partners during both training and inference for object detection. MSBCNet progressively achieves joint cross-modal learning through a multi-stage guided mechanism (coarse perception, modality fusion, fine localization). Our design comprises a video teacher network and audio-visual student network, utilizing visual/audio-visual feature consistency losses and a Cross-modal Feature Aggregation Module (CMFAM) to learn rich video features from unlabeled data. Additionally, a contrastive distribution alignment loss and Salient Audio-Visual Response Module (SAVRM) enhance audio-visual coupling strength. Evaluations on a public multimodal audio-visual detection dataset (87.38% mAP) and self-constructed railway audio-visual detection dataset (81.42% mAP) demonstrate MSBCNet's significant superiority in audio-visual cross-modal object detection, with ablation studies confirming component efficacy.