Related Experiment Video
Updated: May 24, 2026

05:48
Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
Published on: August 9, 2024
Automatic speech recognition in cocktail-party situations: a specific training for separated speech.
Amparo Marti1, Maximo Cobos, Jose J Lopez
1Institute of Telecommunications and Multimedia Applications, Universitat Politècnica de València, 46022, Valencia, Spain. ammargue@iteam.upv.es
The Journal of the Acoustical Society of America
|February 23, 2012
Summary
This study introduces a novel training method to enhance automatic speech recognition (ASR) accuracy in noisy environments. Combining source separation with specialized training significantly improves word recognition for simultaneous speech.
Area of Science:
- Acoustic Signal Processing
- Speech Recognition Technology
- Artificial Intelligence
Background:
- Automatic speech recognition (ASR) systems struggle with accuracy, particularly in adverse acoustic conditions like cocktail-party environments with simultaneous speech.
- Existing source separation methods for simultaneous speech produce artifacts that further degrade ASR performance.
- Human speech recognition capabilities far exceed current ASR system performance in complex acoustic scenarios.
Purpose of the Study:
- To propose a specific training methodology to enhance ASR performance in real-world simultaneous speech scenarios.
- To investigate the combined effectiveness of source separation techniques and the proposed specialized training.
- To evaluate the improvements in ASR accuracy under various acoustic conditions.
Main Methods:
- Development of a specialized training regimen tailored for ASR systems handling simultaneous speech.
- Integration of source separation algorithms to pre-process acoustic signals containing overlapping speech.
- Experimental evaluation of the combined approach across diverse adverse acoustic conditions.
Main Results:
- The proposed specific training significantly improves the percentage of recognized words in simultaneous speech.
- Combining source separation with the specialized training leads to substantial gains in ASR performance.
- Improvements of up to 35% in ASR accuracy were observed under tested acoustical conditions.
Conclusions:
- The developed training strategy effectively addresses the challenges of ASR in simultaneous speech environments.
- The synergistic application of source separation and targeted training offers a viable solution for improving ASR robustness.
- This research presents a significant advancement in achieving more accurate speech recognition in complex acoustic settings.
