Related Experiment Video
Updated: Jun 13, 2025

06:04
Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
Published on: March 24, 2023
351
Comparing human and machine speech recognition in noise with QuickSIN.
Malcolm Slaney1, Matthew B Fitzgerald2
1Center for Computer Research in Music and Acoustics, Stanford University, Stanford, California 94305, USA.
JASA Express Letters
|September 9, 2024
Summary
A new test evaluates automatic speech recognition systems in noise. Modern systems perform similarly to humans, ranging from normal to mildly impaired hearing in noisy conditions.
Area of Science:
- Speech processing
- Human-computer interaction
- Auditory perception
Background:
- Speech recognition systems are crucial for human-computer interaction.
- Evaluating speech recognition in noise is essential for real-world applications.
- Human performance in noise provides a benchmark for system capabilities.
Purpose of the Study:
- To propose a novel test for characterizing automatic speech recognition (ASR) system performance in noise.
- To benchmark modern ASR systems against human performance using the QuickSIN test.
- To establish a standardized metric for evaluating speech-in-noise recognition abilities of ASR.
Main Methods:
- Utilized the QuickSIN (Quick Speech in Noise) test, commonly used in audiology.
- Measured the signal-to-noise ratio (SNR) at which ASR systems achieve 50% keyword recognition.
- Compared ASR performance in noise to established human performance data.
Main Results:
- Modern ASR systems, trained on extensive unsupervised data, were evaluated.
- ASR performance in noise varied, with some systems performing at a normal human level.
- Other systems demonstrated mild impairment in noisy conditions compared to human participants.
Conclusions:
- The proposed test effectively characterizes ASR performance in challenging acoustic environments.
- Modern ASR systems exhibit human-like variability in speech recognition accuracy under noisy conditions.
- Grounding ASR performance metrics to human abilities is vital for developing robust speech technologies.
More Related Videos
Related Concept Videos
Detection of Gross Error: The Q Test
5.7K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
5.7K
Difference from Background: Limit of Detection
6.0K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
6.0K
Perceiving Loudness, Pitch, and Location
202
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
202
Root Mean Square
3.2K
If in an experiment, data values have a probability of being both positive and negative, neither the arithmetic mean, the geometric mean, nor the harmonic mean can be used to calculate the central tendency of the data set. In particular, if the positive and negative values are equally likely, the arithmetic mean is close to zero.
For example, consider the velocity of gas molecules in a container. The gas molecules are moving in different directions, which might impart positive and negative...
For example, consider the velocity of gas molecules in a container. The gas molecules are moving in different directions, which might impart positive and negative...
3.2K
Classification of Signals
420
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
420
Echo
494
The human ear cannot distinguish between two sources of sound if they happen to reach within a specific time interval, typically 0.1 seconds apart. More than this, and they are perceived as separate sources.
Imagine the sound is reflected back to the ears. Assuming that the source is very close to the human, the difference between hearing the two sounds—the emitted sound and the reflected sound—may be more than the minimum time for perceiving distinct sounds. If this is the case,...
Imagine the sound is reflected back to the ears. Assuming that the source is very close to the human, the difference between hearing the two sounds—the emitted sound and the reflected sound—may be more than the minimum time for perceiving distinct sounds. If this is the case,...
494

