Related Experiment Video
Updated: Jan 11, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Effect of masker temporal pattern, spectrum, and presentation level on speech identification
Vijaya Kumar Narne1,2, Saransh Jain3,4, R M Vaibhav5
1Department of Medical Rehabilitation Sciences, College of Applied Medical Sciences, King Khalid University, Abha 61481, Saudi Arabia.
Abstract:
The effects of background spectral and temporal structure and overall level on the masking of speech were assessed for sentences presented in a speech-shaped noise (SSN), harmonic complex tone (HCT) (repetition period = 4.55 ms), and iterated ripple noise (IRN) (delay = 4.55 ms). The speech + noise was presented at 50 and 80 dB SPL using signal-to-noise ratios (SNRs) from 0 to -15 dB in 5-dB steps. The noises had similar spectral envelopes, but the HCT and IRN had spectral dips between peaks corresponding to the harmonic frequencies, and the HCT also had temporal dips. Especially for the SNRs of -5 and -10 dB, speech identification was best for the HCT masker and worst for the SSN, for both overall levels. For the SSN, the SNR required for 50% correct (SNR-50) was higher (worse) at 80 dB than at 50 dB, consistent with poorer frequency selectivity at the higher level. For the HCT, SNR-50 values were lower for the higher level, consistent with a better ability to "listen in the dips" at the higher level. The results indicate that the relative benefit of spectral and temporal dips varies with level.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
06:09P300-Based Brain-Computer Interface Speller Performance Estimation with Classifier-Based Latency Estimation
Published on: September 8, 2023