Related Experiment Videos
FAV-DenoiseNet: An Audio-Visual Speech Enhancement Framework Based on Conditional Flow Matching and Visual Encoding
Xuan Fu1, Lulu Qin1, Weijing Liu1
1School of Computer Science, Jilin Normal University, Siping 136000, China.
Abstract:
Audio-visual speech enhancement aims to recover clean speech by jointly using noisy acoustic signals and synchronized visual cues. Although diffusion-based methods achieve promising restoration performance, their multi-step sampling causes high inference latency and computational cost, limiting real-time deployment. To address this issue, this paper proposes FAV-DenoiseNet, a two-stage framework based on discriminative prior denoising and conditional residual flow matching. The first stage uses a pre-trained discriminative denoising network to suppress dominant noise and provide a structurally stable speech prior. The second stage reformulates enhancement as residual compensation between the first-stage output and the clean speech spectrum instead of directly predicting the entire clean spectrum. A conditional flow-matching network estimates the residual from zero-residual initialization through single-step inference, reducing generative sampling cost. Multi-scale cross-modal attention provides adaptive visual guidance for audio refinement at different resolutions. A residual-controlled fusion strategy preserves the stable structure recovered by the first stage while compensating for residual noise, high-frequency details, and weak speech components. The experimental results show that FAV-DenoiseNet achieves PESQ, ESTOI, and SI-SDR scores of 2.805, 0.775, and 12.480 dB on VoxCeleb2 and 3.157, 0.876, and 13.281 dB on GRID, respectively, with an RTF of 0.086. These results demonstrate that the proposed framework effectively balances enhancement quality, detail restoration, and real-time inference efficiency.
Related Concept Videos
Uniform Depth Channel Flow
Downstream Processing
Deconvolution
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Reconstruction of Signal using Interpolation
Downsampling
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on the metal...