一种基于语音印记细分的双区域语音增强方法
Yang Li1, Wei-Tao Zhang1, Shun-Tian Lou1
1School of Electronic Engineering, Xidian University, Xi'an 710071, China.
概括
这项研究引入了一种新的双区域语音增强模型. 通过将语音细分为不同的区域,该模型可以更好地映射噪音到清洁的语音,显著提高了性能.
科学领域:
- 语音处理 语音处理
- 深度学习是一种深度学习.
- 信号处理 信号处理
背景情况:
- 单通道语音增强通常使用深度学习来将清洁的语音与噪音分开.
- 目前的模型将噪音映射到清洁的演讲,但由于演讲能量分布的变化,这是复杂的.
- 一个单一的模型在高和低语音能量区域中扎着不同的映射要求.
研究的目的:
- 提出一个双区域语音增强模型,以解决单模型方法的局限性.
- 通过单独对待集中和非集中语音能量区域来改善语音增强.
- 为了提高表现并降低语音增强模型的复杂性.
主要方法:
- 开发了一个语音印记细分模型,将噪音语音分为两个不同的区域.
- 单独的深度学习模型被训练为每个识别的语音区域.
- 从双区域模型的输出被合并以重建增强的语音信号.
主要成果:
- 拟议的双区域模型在公共数据集上实现了竞争性的语音增强性能.
- 该方法的性能优于现有的最先进的语音增强技术.
- 废弃性研究证实了双区域方法在改善模型性能方面的有效性.
结论:
- 双区域语音增强模型有效地处理了语音能量分布的独特特征.
- 将不同语音区域的映射任务分开,可以提高整体增强性能.
- 与传统方法相比,这种方法为单通道语音增强提供了更有效的策略.
更多相关视频
相关概念视频
¹³C NMR: Distortionless Enhancement by Polarization Transfer (DEPT)
1.3K
When proton-coupled carbon-13 spectra are simplified by a broadband proton decoupling technique, structural information about the coupled protons is lost. Distortionless enhancement by polarization transfer (DEPT) is a technique that provides information on the number of hydrogens attached to each carbon in a molecule. While the DEPT experiment utilizes complex pulse sequences, the pulse delay and flip angle are specifically manipulated. The resulting signals have different phases depending on...
1.3K
IR Frequency Region: Fingerprint Region
2.1K
IR spectra are divided into two main regions: the diagnostic region and the fingerprint region. The diagnostic region of the spectrum lies above 1500 cm−1. The absorptions resulting from single-bond vibrations of the N–H, C–H, and O–H stretch at higher wavenumbers and appear on the left side of the spectrum. The stretching absorptions of the C≡C and C≡N occur between 2100–2300 cm−1. In contrast, those arising from stretching absorptions of the...
2.1K
Extraction: Advanced Methods
1.3K
Metal ions can be separated from one another by complexation with organic ligands–the chelating agent– to form uncharged chelates. Here, the chelating agent must contain hydrophobic groups and behave as a weak acid, losing a proton to bind with the metal. Since most organic ligands used in this process are insoluble or undergo oxidation in the aqueous phase, the chelating agent is initially added to the organic phase and extracted into the aqueous phase. The metal-ligand complex is...
1.3K
Region of Convergence
1.1K
The z-transform is a powerful mathematical tool used in the analysis of discrete-time signals and systems. It is a crucial tool in the analysis of discrete-time systems, but its convergence is limited to specific values of the complex variable z. This range of values, known as the Region of Convergence (ROC), is fundamental in determining the behavior and stability of a system or signal. The ROC defines the region in the complex plane where the z-transform converges, which can take various...
1.1K
Perceiving Loudness, Pitch, and Location
1.3K
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
1.3K


