Related Experiment Video
Updated: Feb 11, 2026

08:32
Ultrasound Images of the Tongue: A Tutorial for Assessment and Remediation of Speech Sound Errors
Published on: January 3, 2017
23.2K
Noise-robust speech triage.
Anthony L Bartos1, Tomas Cipr2, Douglas J Nelson3
1Suzanne R. Miller Associates, Marriotsville, Maryland 21104, USA.
The Journal of the Acoustical Society of America
|May 3, 2018
Summary
This study enhances speech algorithms for noisy environments by using multiple i-vector models and optimized speech activity detection (SAD). These methods significantly improve speaker identification (SID) and other speech tasks in challenging acoustic conditions.
Area of Science:
- Speech processing
- Signal processing
- Machine learning
Background:
- Conventional speech algorithms struggle in extremely noisy environments.
- Speaker identification (SID) performance degrades significantly with low signal-to-noise ratio (SNR).
- Existing methods often require modifications to adapt to varying noise levels.
Purpose of the Study:
- To improve the performance of speech algorithms in extremely noisy environments without algorithm modification.
- To enhance speaker identification (SID), language identification, gender identification, and diarization.
- To develop robust speech activity detection (SAD) for low-SNR conditions.
Main Methods:
- Applied multiple i-vector algorithms without modification to existing speech processing pipelines.
- Pre-trained multiple speaker identification (SID) models across a range of signal-to-noise-ratio (SNR) levels.
- Implemented two optimized speech activity detection (SAD) algorithms: one for low SNR using voiced speech envelope, another for higher SNR using original speech features.
Main Results:
- Significantly improved classification accuracy (equal error rate) and processing throughput using i-vector algorithms.
- Achieved robust performance in speaker identification, language identification, gender identification, and diarization across all tested SNR levels.
- Demonstrated optimized SID performance when training and testing SNR levels were closely matched.
Conclusions:
- Unmodified conventional speech algorithms, when combined with i-vector techniques and adaptive SAD, can achieve high performance in extremely noisy conditions.
- The proposed approach offers a significant advancement in robust speech processing for real-world noisy environments.
- Optimized SAD is critical for reliable performance of downstream speech tasks when speech is barely audible.
Related Concept Videos
Hearing
57.4K
When we hear a sound, our nervous system is detecting sound waves—pressure waves of mechanical energy traveling through a medium. The frequency of the wave is perceived as pitch, while the amplitude is perceived as loudness.
57.4K
Longitudinal Research
13.5K
Sometimes we want to see how people change over time, as in studies of human development and lifespan. When we test the same group of individuals repeatedly over an extended period of time, we are conducting longitudinal research. Longitudinal research is a research design in which data-gathering is administered repeatedly over an extended period of time. For example, we may survey a group of individuals about their dietary habits at age 20, retest them a decade later at age 30, and then again...
13.5K
Responses to Drought and Flooding
12.1K
Water plays a significant role in the life cycle of plants. However, insufficient or excess of water can be detrimental and pose a serious threat to plants.
12.1K
Self-Discrepancy Theory
18.9K
One influential perspective on what motivates people's behavior is detailed in Tory Higgin's self-discrepancy theory (Higgins, 1987). He proposed that people hold disagreeing internal representations of themselves that lead to different emotional states.
18.9K

