ctPuLSE:靠近的谈话,和基于伪标签的远场,语音增强
1Department of Computer Science and Engineering, Southern University of Science and Technology, Shenzhen 518055, Guangdong, People's Republic of China.
The Journal of the Acoustical Society of America
|October 13, 2025
概括
本研究介绍了ctPuLSE,这是一种用于直接在现实记录上训练神经语音增强模型的新方法. 它使用近距离语音增强来创建伪标签,以提高远距离语音增强的概括性.
科学领域:
- 语音处理 语音处理
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 目前的神经语音增强模型依赖于模拟数据,限制了现实世界的性能.
- 由于缺乏清洁言论监督,直接在真实混合物上进行培训具有挑战性.
研究的目的:
- 开发一种方法,直接在现实记录的混合物上训练语音增强模型.
- 提高远程语音增强模型对现实世界的条件的通用性.
主要方法:
- 建议 ctPuLSE:在模拟数据上训练一个增强模型,以处理真实近距离混合.
- 使用增强的近距离演讲作为伪标签来训练远场增强模型对真实配对混合物的训练.
主要成果:
- ctPuLSE有效地从真正的近距离混合物中生成高质量的伪标签.
- 拟议的方法显著提高了远场语音增强模型在真实数据上的通用性.
结论:
- 近距离对话伪标签提供了一个可行的解决方案,用于对现实世界语音数据的监督培训.
- 在实际场景中,ctPuLSE展示了强大的神经语音增强的有希望的方法.
相关概念视频
Echo
886
The human ear cannot distinguish between two sources of sound if they happen to reach within a specific time interval, typically 0.1 seconds apart. More than this, and they are perceived as separate sources.
Imagine the sound is reflected back to the ears. Assuming that the source is very close to the human, the difference between hearing the two sounds—the emitted sound and the reflected sound—may be more than the minimum time for perceiving distinct sounds. If this is the case,...
Imagine the sound is reflected back to the ears. Assuming that the source is very close to the human, the difference between hearing the two sounds—the emitted sound and the reflected sound—may be more than the minimum time for perceiving distinct sounds. If this is the case,...
886
Perceiving Loudness, Pitch, and Location
935
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
935
Non-Verbal Cues
279
Non-verbal communication extends beyond gestures and facial expressions to include vocal elements known as paralanguage. Paralanguage consists of non-verbal vocal cues such as pitch, loudness, speech rate, pauses, and non-verbal vocalizations like laughter, sighs, and moans. These elements not only accompany speech but also provide critical emotional and contextual information.The Role of Paralanguage in CommunicationParalanguage adds depth to spoken language by conveying emotions and...
279
¹³C NMR: Distortionless Enhancement by Polarization Transfer (DEPT)
1.6K
When proton-coupled carbon-13 spectra are simplified by a broadband proton decoupling technique, structural information about the coupled protons is lost. Distortionless enhancement by polarization transfer (DEPT) is a technique that provides information on the number of hydrogens attached to each carbon in a molecule. While the DEPT experiment utilizes complex pulse sequences, the pulse delay and flip angle are specifically manipulated. The resulting signals have different phases depending on...
1.6K
Interference: Path Lengths
1.9K
Consider two sources of sound, that may or may not be in phase, emitting waves at a single frequency, and consider the frequencies to be the same.
Two special sources may be considered when they are in phase. This can be easily achieved by feeding the two sources from the same source. An example would be synchronizing the two speakers by feeding them with the same source, such as the sound waves produced by a tuning fork. This setup ensures that the two sources have the same frequency and are...
Two special sources may be considered when they are in phase. This can be easily achieved by feeding the two sources from the same source. An example would be synchronizing the two speakers by feeding them with the same source, such as the sound waves produced by a tuning fork. This setup ensures that the two sources have the same frequency and are...
1.9K
Linear Approximation in Frequency Domain
357
Linear systems are characterized by two main properties: superposition and homogeneity. Superposition allows the response to multiple inputs to be the sum of the responses to each individual input. Homogeneity ensures that scaling an input by a scalar results in the response being scaled by the same scalar.
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
357


