Related Experiment Video
Updated: Jan 3, 2026

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
2.0K
Spectral and Temporal Envelope Cues for Human and Automatic Speech Recognition in Noise
Guangxin Hu1, Sarah C Determan1, Yue Dong1
1Biomedical Engineering Department, Saint Louis University, 3007 Lindell Blvd Suite 2007, St Louis, MO, 63103, USA.
Journal of the Association for Research in Otolaryngology : JARO
|November 24, 2019
Summary
Humans and automated speech recognition (ASR) use different strategies in noise. This study found that while temporal cues are vital for humans, ASR relies on spectral analysis, leading to distinct performance differences.
Area of Science:
- Speech processing
- Auditory perception
- Human-computer interaction
Background:
- Speech recognition relies on acoustic features, including spectral and temporal cues.
- Human listeners prioritize temporal envelope, while automated speech recognition (ASR) emphasizes spectral analysis.
- Understanding how noise impacts human and ASR performance is crucial for developing robust systems.
Purpose of the Study:
- To compare sentence recognition scores between humans and ASR (Dragon) under manipulated spectral and temporal cues in various noise conditions.
- To investigate the differential effects of noise types and spectral resolution on human and ASR performance.
- To elucidate the distinct listening strategies employed by humans and ASR in noisy environments.
Main Methods:
- Speech recognition tests were conducted with humans and Dragon ASR software.
- Temporal fine structure was degraded using noise or tone vocoders.
- Spectral information was altered by varying the number of frequency channels.
- Three noise types were used: white noise, time-reversed multi-talker noise, and fake-formant noise.
Main Results:
- White noise more disruptive than fake-formant noise for humans at 20 dB SNR with 4 vocoding channels.
- ASR struggled significantly with 4 vocoding channels even at 20 dB SNR.
- Fake-formant noise most severely impacted ASR, while white noise affected speech segmentation.
- Increasing spectral resolution yielded non-monotonic ASR behavior with white noise.
- ASR performance improved with tone vocoders.
Conclusions:
- Human listeners and ASR systems exhibit fundamentally different processing strategies in noisy conditions.
- Temporal envelope cues are critical for human speech recognition, whereas spectral cues are more important for ASR.
- Noise characteristics significantly influence the performance of both humans and ASR, but in distinct ways.
- Fake-formant noise disrupts ASR by interfering with spectral cues, while white noise impacts speech segmentation.
Related Concept Videos
Perceiving Loudness, Pitch, and Location
860
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
860
Non-Verbal Cues
225
Non-verbal communication extends beyond gestures and facial expressions to include vocal elements known as paralanguage. Paralanguage consists of non-verbal vocal cues such as pitch, loudness, speech rate, pauses, and non-verbal vocalizations like laughter, sighs, and moans. These elements not only accompany speech but also provide critical emotional and contextual information.The Role of Paralanguage in CommunicationParalanguage adds depth to spoken language by conveying emotions and...
225
Perception of Sound Waves
5.4K
The human ear is not equally sensitive to all frequencies in the audible range. It may perceive sound waves with the same pressure but different frequencies as having different loudness. Moreover, the perception of sound waves depends on the health of an individual's ears, which decays with age. The health of one's ears may also be affected by regular exposure to loud noises.
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
5.4K
Hearing
56.3K
When we hear a sound, our nervous system is detecting sound waves—pressure waves of mechanical energy traveling through a medium. The frequency of the wave is perceived as pitch, while the amplitude is perceived as loudness.
56.3K
Auditory Perception
955
The auditory system is essential for sound perception, utilizing various critical structures. When sound waves enter the outer ear, they travel through the ear canal and cause the eardrum to vibrate. These vibrations are then transmitted to the middle ear, where three tiny bones – the malleus, incus, and stapes – amplify the sound. This amplification is crucial, as it ensures that the sound vibrations are strong enough to be conveyed to the inner ear. These vibrations then reach the...
955

