基于格子卷积和对抗训练的光谱网络,用于噪声强大的语音超分辨率
Junkang Yang1, Hongqing Liu1,2, Lu Gan3
1School of Communications and Information Engineering, Chongqing University of Posts and Telecommunications, Chongqing 400065, China.
The Journal of the Acoustical Society of America
|November 12, 2024
概括
这项研究介绍了Super Denoise Net (SDNet),这是一个用于强大的语音超分辨率的新型神经网络. SDNet有效地提高了低分辨率的音频在杂的环境和不同的采样率,优于现有的方法.
科学领域:
- 信号处理 信号处理
- 人工智能的人工智能
背景情况:
- 语音超分辨率将低分辨率的音频增强到高分辨率.
- 现有的模型在与现实世界的噪音和灵活的采样率作斗争.
- 实践应用中的稳定性是一个关键的挑战.
研究的目的:
- 开发一个噪音强大的语音超分辨率模型,适应灵活的输入采样速率.
- 介绍超级Denoise网 (SDNet) 的实用,现实世界的音频增强.
主要方法:
- 设计的超级网 (SDNet) 采用封闭式和格子式卷积块.
- 使用的频率转换块用于长频率依赖模型.
- 在多对手损失培训中使用多级别的区分器.
主要成果:
- 与最先进的模型相比,SDNet表现出卓越的性能.
- 在噪音强大的语音超分辨率方面取得了显著的改进.
- 在多个测试集中验证的有效性.
结论:
- SDNet为现实世界的语音超分辨率挑战提供了强大的解决方案.
- 该模型的设计有效地处理噪音和可变的采样率.
- 这表明实际音频增强技术取得了重大进展.
相关概念视频
Linear Approximation in Frequency Domain
85
Linear systems are characterized by two main properties: superposition and homogeneity. Superposition allows the response to multiple inputs to be the sum of the responses to each individual input. Homogeneity ensures that scaling an input by a scalar results in the response being scaled by the same scalar.
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
85
Deconvolution
133
Deconvolution, also known as inverse filtering, is the process of extracting the impulse response from known input and output signals. This technique is vital in scenarios where the system's characteristics are unknown, and they must be inferred from the observable signals.
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
133
Reducing Line Loss
144
In a three-phase circuit, line loss is an indicator of energy dissipated as heat due to the resistance of transmission lines. To address this, incorporating transformers into the system—a step-up transformer at the source and a step-down transformer at the load—is a strategic solution. Two three-phase transformers are introduced to improve this.
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
144
Perceiving Loudness, Pitch, and Location
195
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
195
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Downsampling
131
When considering a sampled sequence with zero values between sampling instants, one can replace it by taking every N-th value of the sequence. At these integer multiples of N, the original and sampled sequences coincide. This process, known as decimation, involves extracting every N-th sample from a sequence, thereby creating a more efficient sequence.
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
131


