Related Experiment Video
Updated: Nov 11, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
D-MONA: A dilated mixed-order non-local attention network for speaker and language recognition
Xiaoxiao Miao1, Ian McLoughlin2, Wenchao Wang1
1Key Laboratory of Speech Acoustics and Content Understanding, Institute of Acoustics, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences, Beijing, China.
This study introduces a novel dilated mixed-order non-local attention network (D-MONA) for improved speaker and language recognition. D-MONA enhances feature extraction by analyzing multi-order speech information across wider contexts, outperforming existing methods.
Area of Science:
- Speech processing and machine learning
- Deep learning for audio analysis
- Speaker and language recognition technologies
Background:
- Convolutional neural network (CNN) models with attention mechanisms are common for speaker and language recognition (SR/LR).
- Existing attention methods struggle with complex information and multi-scale, long-range speech feature interactions.
- There is a need for advanced attention mechanisms to capture richer speech feature representations.
Purpose of the Study:
- To propose a novel attention mechanism that addresses limitations in current SR/LR models.
- To enhance the extraction of fine-grained, multi-order speech features and long-range interactions.
- To improve the performance of speaker and language recognition systems.
Main Methods:
- Introduction of mixed-order attention (MOA) for detailed frame-level speech features.
- Integration of non-local attention (NLA) and a dilated residual structure.
- Development of the dilated mixed-order non-local attention network (D-MONA).
Main Results:
- D-MONA significantly improves SR performance on Voxceleb1 (29%) and CN-celeb (15%) compared to ResNet-34.
- For LR tasks, D-MONA shows substantial gains over ResNet-34, especially for shorter utterances (21% for 3s, 59% for 10s, 67% for 30s).
- The proposed D-MONA outperforms the state-of-the-art DBF-DNN x-vector system across all tested utterance lengths.
Conclusions:
- The D-MONA model effectively captures multi-scale, long-range speech feature interactions.
- This novel network architecture offers superior performance in both speaker and language recognition tasks.
- D-MONA represents a significant advancement in deep learning for SR/LR applications.

