DSF-Net: Dual-strategy fusion for efficient audio-visual sound event localization and detection

Rendong Pi1, Yingchao Zhang2, Wei Rao3

  • 1Department of Mechanical Engineering, The Hong Kong Polytechnic University, Kowloon, Hong Kong, China; College of Computing and Data Science, Nanyang Technological University, Singapore.

Summary

DSF-Net enhances audio-visual sound event localization and detection (AVSELD) using a novel dual-strategy fusion approach. This method achieves state-of-the-art performance efficiently by employing a state-space model backbone.

Related Concept Videos