Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Videos

Exploiting audio-visual modalities in videos: Object detection via multi-stage bilateral coupling network.

Qifeng Liu1, Zujun Yu2, Liqiang Zhu2

  • 1State Key Laboratory of Advanced Rail Autonomous Operation, Beijing Jiaotong University, Beijing, 100044, Beijing, China; School of Mechanical, Electronic and Control Engineering, Beijing Jiaotong University, Beijing, 100044, Beijing, China.

Neural Networks : the Official Journal of the International Neural Network Society
|June 30, 2026
PubMed
Summary

Related Concept Videos

Multi-input and Multi-variable systems01:22

Multi-input and Multi-variable systems

Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

From health to periodontitis: dynamic changes in the subgingival microbiome and their association with systemic inflammation levels.

Frontiers in cellular and infection microbiology·2026
Same author

Identification of a novel ABO*A1.01 allele with c.562C>T (p.Arg188Cys) mutation associated with A<sub>el</sub> phenotype in a Chinese individual.

Transfusion·2026
Same author

The Transcription Factors HbWRKY29 and HbPTI5 cooperatively enhance rubber tree resistance to powdery mildew.

Molecular plant pathology·2026
Same author

Multimodal deep learning model for AI-based functional prognostic risk stratification in patients undergoing radical nephrectomy.

Nature communications·2026
Same author

Comparative Evaluation of Buffy Coat-Derived and Apheresis Platelet-Rich Plasma in the Treatment of Androgenetic Alopecia: Laboratory and Clinical Insights.

Clinical, cosmetic and investigational dermatology·2026
Same author

Pediatric gastric MALT lymphoma: an antibiotic-treatable lymphoma, a case report and review of the literature.

BMC pediatrics·2026

The Multi-Stage Bilateral Coupling Network (MSBCNet) advances audio-visual object detection by treating audio and visual data equally during training and inference. This novel approach enables the detection of both sounding and non-sounding objects, improving comprehensive perception.

Area of Science:

  • Computer Vision
  • Machine Learning
  • Signal Processing

Background:

  • Existing audio-visual object detection methods often use asymmetric fusion, neglecting one modality during inference.
  • Current methods primarily focus on sound source localization, limiting detection to sounding objects.
  • This prevents the full exploitation of integrated audio-visual data for comprehensive perception.

Purpose of the Study:

  • To develop an object detection method that effectively utilizes both audio and visual modalities as co-equal partners.
  • To enable the detection of both sounding and non-sounding objects in video data.
  • To introduce the Multi-Stage Bilateral Coupling Network (MSBCNet) for holistic audio-visual learning.

Main Methods:

  • Proposed the Multi-Stage Bilateral Coupling Network (MSBCNet) with a multi-stage guided mechanism (coarse perception, modality fusion, fine localization).
Keywords:
Audio-visual object detectionFeature couplingModality fusionMulti-stage guided

Related Experiment Videos

  • Employed a video teacher and audio-visual student network design, utilizing feature consistency losses and a Cross-modal Feature Aggregation Module (CMFAM).
  • Incorporated contrastive distribution alignment loss and a Salient Audio-Visual Response Module (SAVRM) to strengthen audio-visual coupling.
  • Main Results:

    • MSBCNet achieved 87.38% mAP on a public multimodal dataset and 81.42% mAP on a railway dataset.
    • Demonstrated significant superiority in audio-visual cross-modal object detection compared to existing methods.
    • Ablation studies confirmed the efficacy of individual components within the MSBCNet framework.

    Conclusions:

    • MSBCNet is the first framework to holistically leverage audio and visual modalities as co-equal partners for object detection during both training and inference.
    • The proposed method effectively addresses the limitations of existing approaches, enabling detection of both sounding and non-sounding objects.
    • MSBCNet significantly advances the field of audio-visual cross-modal object detection.