Related Experiment Video
Updated: Aug 8, 2026

07:52
An Automated System for Sound Localization Testing in Hearing-Impaired Listeners
Published on: March 13, 2026
DSF-Net: Dual-strategy fusion for efficient audio-visual sound event localization and detection
Rendong Pi1, Yingchao Zhang2, Wei Rao3
1Department of Mechanical Engineering, The Hong Kong Polytechnic University, Kowloon, Hong Kong, China; College of Computing and Data Science, Nanyang Technological University, Singapore.
Summary
DSF-Net enhances audio-visual sound event localization and detection (AVSELD) using a novel dual-strategy fusion approach. This method achieves state-of-the-art performance efficiently, overcoming limitations of existing models.
Area of Science:
- Computer Vision
- Machine Learning
- Signal Processing
Background:
- Audio-visual sound event localization and detection (AVSELD) traditionally uses Convolutional Neural Networks (CNNs).
- CNNs have limited receptive fields, hindering capture of broader contextual information.
- Transformer models capture global context but suffer from quadratic computational complexity with long sequences.
Purpose of the Study:
- To introduce DSF-Net, a novel neural network for AVSELD.
- To address computational bottlenecks in processing long-range dependencies in AVSELD.
- To achieve robust and computationally efficient multi-modal comprehension for AVSELD.
Main Methods:
- Developed DSF-Net utilizing an efficient state-space model backbone for linear complexity.
- Introduced a dual-strategy fusion approach: Adaptive Frequency Fusion and Audio-aware Aggregation modules.
- Integrated these strategies within a progressive fusion framework for enhanced feature learning.
Main Results:
- DSF-Net demonstrated state-of-the-art performance on the STARSS2023 dataset.
- The dual-strategy fusion approach significantly outperformed existing AVSELD methods.
- The proposed model offers robust and computationally efficient multi-modal comprehension.
Conclusions:
- DSF-Net effectively overcomes the limitations of previous AVSELD approaches.
- The novel fusion strategies provide significant improvements in performance and efficiency.
- Publicly available source code facilitates further research and development in AVSELD.