Related Experiment Video
Updated: May 29, 2026

09:25
Pupillometry to Assess Auditory Sensation in Guinea Pigs
Published on: January 6, 2023
Predicting speech intelligibility based on the signal-to-noise envelope power ratio after modulation-frequency
1Centre for Applied Hearing Research, Department of Electrical Engineering, Technical University of Denmark, DK-2800 Kgs. Lyngby, Denmark.
The Journal of the Acoustical Society of America
|September 8, 2011
Summary
A new speech intelligibility model predicts how well people understand processed noisy speech. It uses the speech-to-noise envelope power ratio (SNRenv) and accurately predicts intelligibility across various noise conditions.
Area of Science:
- Acoustics
- Speech Processing
- Psychoacoustics
Background:
- Accurate prediction of speech intelligibility in noisy conditions is crucial for developing effective hearing aids and communication systems.
- Existing models, like the speech transmission index, may not fully capture intelligibility in complex noise scenarios such as spectral subtraction.
- Understanding the impact of noise characteristics on speech perception is an ongoing area of research.
Purpose of the Study:
- To propose and validate a novel model for predicting the intelligibility of processed noisy speech.
- To assess the model's performance across different noise types, including stationary noise, reverberation, and spectral subtraction.
- To investigate the underlying mechanisms contributing to intelligibility changes, particularly in spectral subtraction conditions.
Main Methods:
- Developed a speech-based envelope power spectrum model, drawing parallels with existing modulation detection models.
- Estimated the speech-to-noise envelope power ratio (SNRenv) at the output of a modulation filterbank.
- Related the SNRenv metric to speech intelligibility using an ideal observer framework.
- Compared model predictions with empirical intelligibility data for speech in various noise conditions.
Main Results:
- The proposed model demonstrated good agreement with intelligibility data for speech in stationary noise, reverberation, and spectral subtraction.
- Analysis revealed that for spectral subtraction, intelligibility reduction was linked to the estimated noise envelope power exceeding speech envelope power.
- The model successfully predicted intelligibility where the classical speech transmission index failed, particularly in spectral subtraction scenarios.
Conclusions:
- The speech-to-noise envelope power ratio (SNRenv) derived from a modulation filterbank is a robust predictor of speech intelligibility.
- The proposed model offers improved accuracy over traditional methods for processed noisy speech.
- Modulation-based signal-to-noise ratio measures are key to understanding and predicting speech intelligibility in challenging acoustic environments.