D-MONA: A dilated mixed-order non-local attention network for speaker and language recognition

Xiaoxiao Miao1, Ian McLoughlin2, Wenchao Wang1

  • 1Key Laboratory of Speech Acoustics and Content Understanding, Institute of Acoustics, Chinese Academy of Sciences, Beijing, China; University of Chinese Academy of Sciences, Beijing, China.

Summary

This study introduces a novel dilated mixed-order non-local attention network (D-MONA) for improved speaker and language recognition. D-MONA enhances feature extraction by analyzing multi-order speech information across wider contexts, outperforming existing methods.

Related Concept Videos