基于集群的对比对比损失对于噪音-强大的语音识别
Geon Woo Lee1, Hong Kook Kim1,2,3
1AI Graduate School, Gwangju Institute of Science and Technology, Gwangju 61005, Republic of Korea.
Sensors (Basel, Switzerland)
|April 27, 2024
概括
本研究介绍了一种用于语音增强 (SE) 和自动语音识别 (ASR) 的联合训练方法,使用一种新的基于集群的对对比 (CBPC) 损失. 这种方法提高了语音质量,并降低了在噪音条件下文字错误率 (WER).
科学领域:
- 人工智能的人工智能
- 语音处理 语音处理
- 机器学习 机器学习
背景情况:
- 语音增强 (SE) 和自动语音识别 (ASR) 模型的联合培训对于提高语音处理性能至关重要.
- 从ASR到SE模型中利用语言信息可以增强SE的能力.
- 现有的方法可能无法完全捕捉语言细微差别,或者可能过度适应杂的数据.
研究的目的:
- 为SE和ASR模型提出一种新的联合培训方法.
- 为有效的语言信息传输引入声学标记器和基于集群的对对比 (CBPC) 损失函数.
- 提高语音增强质量,降低噪音环境中的文字错误率 (WER).
主要方法:
- 开发了一个管道,将SE和ASR模型与声学标记器集成在一起.
- 声学标记器使用来自ASR编码器输出的K-means集群生成伪标签.
- 一个新的基于集群的对比对比 (CBPC) 损失函数,结合infoNCE,被提出用于自我监督学习以传输语言信息.
主要成果:
- 与CBPC损失一起提出的联合培训方法与传统方法相比,实现了较低的文字错误率 (WER).
- 语音质量得分明显高于独立的SE模型和使用传统联合方法培训的SE模型.
- 结合CBPC和infoNCE损失,在减少WER和提高语音质量方面表现出有效性.
结论:
- 拟议的CBPC损失功能有效地将语言信息从ASR转移到SE模型,在联合培训框架内.
- 这种方法导致了优越的语音增强性能和在噪音条件下更准确的语音识别.
- 该方法为推进强大的语音处理系统提供了一个有希望的方向.
更多相关视频
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
443
06:04Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
Published on: March 24, 2023
380
相关概念视频
Reducing Line Loss
151
In a three-phase circuit, line loss is an indicator of energy dissipated as heat due to the resistance of transmission lines. To address this, incorporating transformers into the system—a step-up transformer at the source and a step-down transformer at the load—is a strategic solution. Two three-phase transformers are introduced to improve this.
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
151
Difference from Background: Limit of Detection
6.4K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
6.4K
Wilcoxon Signed-Ranks Test for Matched Pairs
121
The Wilcoxon signed-rank test for matched pairs evaluates the null hypothesis by combining the ranks of differences with their signs. It essentially tests whether the median of the differences in a population of matched pairs is zero. Since the test incorporates more information than the sign test, it generally yields more trustable conclusions. This test also does not require the data to follow a normal distribution, but two conditions must be met for it to be applicable: (1) the data must...
121
Line Loss
245
The different configurations of source-load connections include wye (star) and delta connections. The relationship between line and phase voltages and currents varies depending on the configuration. When the source is supplying power, it is transmitted through the wires to the load, and during this transmission, some power is absorbed by the wires, leading to line loss.
Line loss impacts power delivery efficiency in a balanced three-phase circuit. The symmetry in such a circuit simplifies the...
Line loss impacts power delivery efficiency in a balanced three-phase circuit. The symmetry in such a circuit simplifies the...
245
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
