将人类和机器在噪音中的语音识别与QuickSIN进行比较
Malcolm Slaney1, Matthew B Fitzgerald2
1Center for Computer Research in Music and Acoustics, Stanford University, Stanford, California 94305, USA.
JASA express letters
|September 9, 2024
概括
一项新的测试评估了噪音中的自动语音识别系统. 现代系统的性能与人类相似,在噪音条件下从正常到轻微受损的听力范围.
科学领域:
- 语音处理 语音处理
- 人与计算机的互动.
- 听觉感知是一种听觉感知.
背景情况:
- 语音识别系统对于人机交互至关重要.
- 在噪音中评估语音识别对于现实世界的应用是必不可少的.
- 人类在噪声中的表现为系统能力提供了一个基准.
研究的目的:
- 提出一种新型测试,用于描述自动语音识别 (ASR) 系统在噪音中的性能.
- 使用QuickSIN测试对现代ASR系统与人类性能进行基准测试.
- 建立一个标准化的指标来评估ASR的语音噪音识别能力.
主要方法:
- 使用了QuickSIN (噪音中的快速言语) 测试,通常用于听力学.
- 测量了ASR系统实现50%关键词识别的信号噪声比 (SNR).
- 将ASR在噪声中的性能与已确定的人类性能数据进行比较.
主要成果:
- 现代的ASR系统,训练在广泛的无监督数据,被评估.
- 在噪音中的ASR性能各不相同,有些系统在正常人类水平上表现得很好.
- 与人类参与者相比,其他系统在噪音条件下表现出轻微损伤.
结论:
- 拟议的测试有效地描述了ASR在具有挑战性的声学环境中的性能.
- 现代ASR系统在杂的环境下表现出语音识别准确度的类似人类的变化.
- 将ASR性能指标与人类能力结合起来,对于开发强大的语音技术至关重要.
更多相关视频
相关概念视频
Detection of Gross Error: The Q Test
5.7K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
5.7K
Difference from Background: Limit of Detection
6.0K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
6.0K
Perceiving Loudness, Pitch, and Location
202
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
202
Root Mean Square
3.2K
If in an experiment, data values have a probability of being both positive and negative, neither the arithmetic mean, the geometric mean, nor the harmonic mean can be used to calculate the central tendency of the data set. In particular, if the positive and negative values are equally likely, the arithmetic mean is close to zero.
For example, consider the velocity of gas molecules in a container. The gas molecules are moving in different directions, which might impart positive and negative...
For example, consider the velocity of gas molecules in a container. The gas molecules are moving in different directions, which might impart positive and negative...
3.2K
Classification of Signals
420
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
420
Echo
494
The human ear cannot distinguish between two sources of sound if they happen to reach within a specific time interval, typically 0.1 seconds apart. More than this, and they are perceived as separate sources.
Imagine the sound is reflected back to the ears. Assuming that the source is very close to the human, the difference between hearing the two sounds—the emitted sound and the reflected sound—may be more than the minimum time for perceiving distinct sounds. If this is the case,...
Imagine the sound is reflected back to the ears. Assuming that the source is very close to the human, the difference between hearing the two sounds—the emitted sound and the reflected sound—may be more than the minimum time for perceiving distinct sounds. If this is the case,...
494


