Related Experiment Videos
FAV-DenoiseNet: An Audio-Visual Speech Enhancement Framework Based on Conditional Flow Matching and Visual Encoding
Xuan Fu1, Lulu Qin1, Weijing Liu1
1School of Computer Science, Jilin Normal University, Siping 136000, China.
Sensors (Basel, Switzerland)
|July 15, 2026
Summary
This study introduces FAV-DenoiseNet, a novel two-stage framework for audio-visual speech enhancement. It significantly reduces latency for real-time applications by using discriminative denoising and flow matching for efficient speech recovery.
Area of Science:
- Speech processing
- Artificial intelligence
- Signal processing
Background:
- Diffusion-based methods offer high performance in audio-visual speech enhancement but suffer from high latency.
- Real-time deployment of speech enhancement systems is hindered by computational costs.
- Existing methods struggle to balance enhancement quality with inference speed.
Purpose of the Study:
- To develop an efficient audio-visual speech enhancement framework (FAV-DenoiseNet) that overcomes the latency limitations of diffusion models.
- To improve the quality of enhanced speech by effectively integrating visual cues.
- To achieve real-time performance without compromising restoration accuracy.
Main Methods:
- A two-stage framework combining discriminative prior denoising and conditional residual flow matching.
- The first stage suppresses noise and provides a stable speech prior.
- The second stage estimates the residual using single-step flow matching with multi-scale cross-modal attention and a residual-controlled fusion strategy.
Main Results:
- FAV-DenoiseNet achieves state-of-the-art performance on benchmark datasets (VoxCeleb2, GRID) with high PESQ, ESTOI, and SI-SDR scores.
- The proposed method demonstrates a low real-time factor (RTF) of 0.086, enabling efficient inference.
- The framework effectively balances speech enhancement quality, detail restoration, and real-time processing.
Conclusions:
- FAV-DenoiseNet successfully addresses the latency and computational cost issues of diffusion-based speech enhancement.
- The proposed two-stage approach with residual compensation and cross-modal attention provides superior audio-visual speech enhancement.
- The framework offers a promising solution for real-time audio-visual speech enhancement applications.
Related Concept Videos
Uniform Depth Channel Flow
Uniform depth channel flow keeps fluid depth consistent along channels such as irrigation canals. In natural channels, such as rivers, approximate uniform flow is often assumed. This condition occurs when the channel’s bottom slope matches the energy slope, balancing potential energy lost from gravity with head loss due to shear stress. This balance prevents depth changes along the channel length, resulting in a steady, uniform flow.Uniform flow in open channels with a constant cross-section...
Downstream Processing
Downstream processing begins once fermentation is complete and involves a series of steps to recover and purify products such as acids, vitamins, antibiotics, or proteins.Cell HarvestingFor example, for intracellular protein-based products, the first step is harvesting the cells. This is typically achieved using centrifugation or filtration to separate the cells from the liquid phase.Cell Disruption for Intracellular ProductsIf the target product is intracellular, the harvested cells must be...
Deconvolution
Deconvolution, also known as inverse filtering, is the process of extracting the impulse response from known input and output signals. This technique is vital in scenarios where the system's characteristics are unknown, and they must be inferred from the observable signals.
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Reconstruction of Signal using Interpolation
Signal processing techniques are essential for accurately converting continuous signals to digital formats and vice versa. When a continuous signal is sampled with a period T, the resulting sampled signal exhibits replicas of the original spectrum in the frequency domain, spaced at intervals equal to the sampling frequency. To handle this sampled signal, a zero-order hold method can be applied, which creates a piecewise constant signal by retaining each sample's value until the next sampling...
Downsampling
When considering a sampled sequence with zero values between sampling instants, one can replace it by taking every N-th value of the sequence. At these integer multiples of N, the original and sampled sequences coincide. This process, known as decimation, involves extracting every N-th sample from a sequence, thereby creating a more efficient sequence.
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
Masking and Demasking Agents
EDTA titrations may necessitate masking and demasking agents to temporarily protect a particular metal ion in a mixture from the EDTA reaction. These agents facilitate the sequential analysis of the metal ions by forming stable complexes with some—but not all—metal ions during certain steps.
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on the metal...
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on the metal...