Related Experiment Video
Updated: May 11, 2025

06:04
Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
Published on: March 24, 2023
311
Evaluating synthesized speech intelligibility in noise
Ye Yang1, Dathan Nguyen1, Katherine Chen1
1Department of Biomedical Engineering, University of California, Irvine, California 92697, USA.
JASA Express Letters
|April 17, 2025
Summary
Synthesized voices, like human speech, can be intelligible in noisy conditions. Some machine-generated voices even outperform human speech intelligibility and are better recognized by automatic speech recognition systems.
Area of Science:
- Speech processing
- Human-computer interaction
- Acoustic phonetics
Background:
- Humans adapt speech for clarity in noise.
- Advancing speech synthesis technology enables machines to generate intelligible speech.
- Evaluating synthesized speech intelligibility in noise is crucial for human-computer interaction.
Purpose of the Study:
- To assess the subjective and objective intelligibility of synthesized speech in speech-shaped noise.
- To compare the intelligibility of synthesized voices from major platforms against human speech.
- To evaluate the performance of automatic speech recognition systems with synthesized speech.
Main Methods:
- Subjective intelligibility tests involving human listeners.
- Objective intelligibility measurements.
- Evaluation of synthesized speech from three major speech synthesis platforms.
- Testing with speech-shaped noise.
Main Results:
- Synthesized voices demonstrated an intelligibility range comparable to human voices.
- Certain synthesized voices exhibited higher intelligibility than human voices.
- Two modern automatic speech recognition systems achieved 10% higher word recognition accuracy compared to human listeners.
Conclusions:
- Speech synthesis technology can produce voices highly intelligible in noisy environments.
- Synthesized voices show potential to match or exceed human speech intelligibility.
- Automatic speech recognition systems perform exceptionally well with intelligible synthesized speech.
Related Concept Videos
Sound Intensity Level
4.1K
Humans perceive sound by hearing. The human ear helps sound waves reach the brain, which then interprets the waves and creates the perception of hearing. The loudness of the environment in which a person is located determines whether they can distinguish between different sound sources.
The human ear can perceive an extensive range of sound intensity, necessitating the use of the logarithmic scale to define a physical quantity—the intensity level. It is a ratio of two intensities and...
The human ear can perceive an extensive range of sound intensity, necessitating the use of the logarithmic scale to define a physical quantity—the intensity level. It is a ratio of two intensities and...
4.1K
Perceiving Loudness, Pitch, and Location
161
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
161
Sound Intensity
4.0K
The loudness of a sound source is related to how energetically the source is vibrating, consequently making the molecules of the propagation medium vibrate. To measure the loudness of a source, the physical quantity of interest is the intensity. This is defined as the energy emitted per unit of time per unit of area perpendicular to the sound wave's propagation direction. Since the total energy is greater if the source vibrates for a longer duration and over a larger area, dividing the...
4.0K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Aliasing
100
Accurate signal sampling and reconstruction are crucial in various signal-processing applications. A time-domain signal's spectrum can be revealed using its Fourier transform. When this signal is sampled at a specific frequency, it results in multiple scaled replicas of the original spectrum in the frequency domain. The spacing of these replicas is determined by the sampling frequency.
If the sampling frequency is below the Nyquist rate, these replicas overlap, preventing the original...
If the sampling frequency is below the Nyquist rate, these replicas overlap, preventing the original...
100

