Related Experiment Video
Updated: Sep 16, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.6K
An enhanced deep learning approach for speaker diarization using TitaNet, MarbelNet and time delay network
Muzamil Ahmed1, Riad Alharbey2, Ali Daud3
1Department of Computer Science, Namal University, Mianwali, 42210, Pakistan.
Scientific Reports
|July 8, 2025
Summary
This study introduces the Neuro-TM Diarizer, a deep learning framework enhancing speaker diarization accuracy. It significantly outperforms traditional methods in complex acoustic environments, improving performance on key metrics.
Area of Science:
- Speech Processing and Signal Analysis
- Machine Learning and Artificial Intelligence
- Computational Linguistics
Background:
- Speaker diarization is crucial for speech transcription, AI, and audio analysis, but traditional methods struggle with noise and overlapping speech.
- Existing clustering-based approaches exhibit limitations in accuracy and robustness, particularly in challenging acoustic conditions.
Purpose of the Study:
- To introduce a novel deep learning framework, the Neuro-TM Diarizer, for improved speaker diarization.
- To enhance performance in complex acoustic environments by integrating noise reduction, adaptive beamforming, and neural diarization.
Main Methods:
- Developed a multimodal deep learning framework combining Marble-Net for voice activity detection and Tita-Net for speaker embeddings.
- Employed time-delay neural networks for neural diarization and speaker identification.
- Evaluated the Neuro-TM Diarizer on VoxConverse and VoxCeleb datasets, comparing it against clustering-based methods.
Main Results:
- The Neuro-TM Diarizer achieved low Diarization Error Rates (DER) of 6.89% on VoxConverse and 6.93% on VoxCeleb.
- Demonstrated significant improvements over clustering-based approaches, reducing DER by 12.60% on VoxConverse and 14.01% on VoxCeleb.
- The framework effectively handles noise and overlapping speech, outperforming traditional methods in accuracy and robustness.
Conclusions:
- The proposed Neuro-TM Diarizer offers a substantial advancement in speaker diarization technology.
- This deep learning framework provides a robust solution for real-world applications including speech transcription, speaker authentication, and audio archiving.
- The integration of noise reduction and advanced neural network architectures leads to superior performance in complex acoustic scenarios.

