Related Experiment Video
Updated: Jan 13, 2026

Sound Source Localization Testing in Single-sided Deafness Following Bone Conduction Intervention
Published on: December 20, 2024
MSFDnet: A Multi-Scale Feature Dual-Layer Fusion Model for Sound Event Localization and Detection
Yi Chen1, Zhenyu Huang2, Liang Lei3
1School of Big Data and Information Industry, Chongqing City Management College, No. 151, Daxuecheng South Second Road, Shapingba District, Chongqing 401331, China.
Abstract:
The task of Sound Event Localization and Detection (SELD) aims to simultaneously address sound event recognition and spatial localization. However, existing SELD methods face limitations in long-duration dynamic audio scenarios, as they do not fully leverage the complementarity between multi-task features and lack depth in feature extraction, leading to restricted system performance. To address these issues, we propose a novel SELD model-MSDFnet. By introducing a Multi-Scale Feature Aggregation (MSFA) module and a Dual-Layer Feature Fusion strategy (DLFF), MSDFnet captures rich spatial features at multiple scales and establishes a stronger complementary relationship between SED and DOA features, thereby enhancing detection and localization accuracy. On the DCASE2020 Task 3 dataset, our model achieved scores of 0.319, 76%, 10.2°, 82.4%, and 0.198 in ER20,F20, LEcd, LRcd, and SELDscore metrics, respectively. Experimental results demonstrate that MSDFnet performs excellently in complex audio scenarios. Additionally, ablation studies further confirm the effectiveness of the MSFA and DLFF modules in enhancing SELD task performance.

