Related Experiment Video
Updated: Aug 19, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Neural speech enhancement with unsupervised pre-training and mixture training
Xiang Hao1, Chenglin Xu2, Lei Xie1
1Audio, Speech and Langauge Processing Group, School of Computer Science, Northwestern Polytechnical University, Xi'an, China.
This study introduces a novel unsupervised pre-training and mixture training algorithm for speech enhancement. The method effectively leverages unpaired data, improving performance over existing supervised and unsupervised approaches.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Signal Processing
Background:
- Supervised speech enhancement relies heavily on large paired noisy and clean speech datasets, which are difficult to obtain in real-world scenarios.
- Simulated data used in supervised methods often leads to a performance gap when deployed due to domain mismatch with in-the-wild data.
- Unsupervised methods avoid paired data but typically underperform compared to supervised techniques.
Purpose of the Study:
- To develop an effective speech enhancement method that overcomes the limitations of both supervised and unsupervised approaches.
- To leverage the strengths of both supervised and unsupervised learning for improved speech enhancement performance.
- To address the data scarcity and domain mismatch issues in real-world speech enhancement applications.
Main Methods:
- Proposes an unsupervised pre-training phase using large volumes of unpaired noisy and clean speech data.
- Introduces a mixture training phase that utilizes in-the-wild noisy data and a small amount of simulated paired data.
- Optimizes a pre-trained model through the combined training strategy.
Main Results:
- The proposed method demonstrates superior performance compared to current state-of-the-art supervised and unsupervised speech enhancement techniques.
- Achieves better performance consistency by mitigating the mismatch between simulated and real-world data.
- Effectively utilizes large-scale unpaired data for initial model training.
Conclusions:
- The unsupervised pre-training and mixture training algorithm offers a robust solution for speech enhancement.
- This hybrid approach successfully bridges the gap between supervised and unsupervised learning methods.
- The findings suggest a promising direction for developing more practical and high-performing speech enhancement systems.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
06:04Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
Published on: March 24, 2023