Related Experiment Video
Updated: Aug 8, 2026

07:52
An Automated System for Sound Localization Testing in Hearing-Impaired Listeners
Published on: March 13, 2026
DSF-Net: Dual-strategy fusion for efficient audio-visual sound event localization and detection
Rendong Pi1, Yingchao Zhang2, Wei Rao3
1Department of Mechanical Engineering, The Hong Kong Polytechnic University, Kowloon, Hong Kong, China; College of Computing and Data Science, Nanyang Technological University, Singapore.
Summary
DSF-Net enhances audio-visual sound event localization and detection (AVSELD) using a novel dual-strategy fusion approach. This method achieves state-of-the-art performance efficiently by employing a state-space model backbone.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Signal Processing
Background:
- Audio-visual sound event localization and detection (AVSELD) traditionally uses Convolutional Neural Networks (CNNs), limited by receptive fields.
- Transformer models capture global context but suffer from quadratic computational complexity with long sequences.
- Existing methods struggle with efficient and comprehensive multi-modal feature fusion for AVSELD.
Purpose of the Study:
- To introduce DSF-Net, a novel neural network for efficient and robust AVSELD.
- To overcome the computational limitations of Transformer models in AVSELD.
- To enhance multi-modal comprehension through a dual-strategy fusion approach.
Main Methods:
- Developed DSF-Net utilizing an efficient state-space model backbone for linear computational complexity.
- Implemented an Adaptive Frequency Fusion module for frequency domain feature alignment and integration.
- Introduced an Audio-aware Aggregation module for advanced feature integration, considering inter-modal consistency.
Main Results:
- DSF-Net demonstrated state-of-the-art performance on the STARSS2023 dataset.
- The proposed dual-strategy fusion approach significantly outperformed existing AVSELD methods.
- The model achieved robust and computationally efficient multi-modal comprehension.
Conclusions:
- DSF-Net offers a computationally efficient and effective solution for AVSELD.
- The dual-strategy fusion approach is crucial for advancing multi-modal audio-visual understanding.
- The public availability of source code facilitates further research and development in AVSELD.