Related Experiment Video
Updated: Sep 27, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Speech Enhancement by Multiple Propagation through the Same Neural Network
Tomasz Grzywalski1, Szymon Drgas1
1Institute of Automatic Control and Robotics, Poznan University of Technology, 60-965 Poznan, Poland.
Repeating speech enhancement processing multiple times improves clarity. This study shows that up to five iterations enhance speech intelligibility, with diminishing gains but notable improvements over single passes, especially for U-Net and Transformer-Net architectures.
Area of Science:
- * Digital Signal Processing
- * Machine Learning
- * Audio Engineering
Background:
- * Monaural speech enhancement aims to improve speech clarity by removing background noise.
- * Deep neural networks are the current state-of-the-art for speech enhancement.
- * Previous research showed multi-forward-pass enhancement with U-Net improved results.
Purpose of the Study:
- * To investigate the effectiveness of multi-forward-pass speech enhancement on different neural network architectures.
- * To determine the performance limits of multi-forward-pass enhancement by testing up to five iterations.
- * To compare the benefits of multi-forward-pass enhancement across U-Net, ResBLSTM, and Transformer-Net models.
Main Methods:
- * Implemented multi-forward-pass speech enhancement for ResBLSTM and Transformer-Net architectures.
- * Tested enhancement iterations from two up to five.
- * Evaluated performance using SI-SDR, STOI, and PESQ metrics on WSJ0, Noisex-92, and DCASE datasets.
Main Results:
- * Multi-forward-pass speech enhancement up to five iterations consistently improved speech intelligibility.
- * Performance gains diminished with each additional iteration.
- * Five iterations yielded an additional 0.6 dB SI-SDR and 4% STOI gain compared to two iterations.
- * U-Net and Transformer-Net architectures benefited more from multi-forward passes than ResBLSTM.
Conclusions:
- * Multi-forward-pass speech enhancement is effective across various architectures, extending beyond U-Net.
- * While gains decrease with more iterations, five passes offer significant improvements over fewer passes.
- * Architecture choice impacts the effectiveness of multi-forward-pass enhancement, with U-Net and Transformer-Net showing greater advantages.
Related Concept Videos
Propagation of Action Potentials
Neurons (nerve cells) have a resting membrane potential, with a slightly negative charge inside compared to outside. This is maintained by ion channels, such as sodium (Na+) and potassium (K+) channels, which control the flow of ions. When a stimulus, like a touch or a signal from another neuron, triggers the neuron, sodium channels open, allowing sodium ions to...
Neural Circuits
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...
Auditory Pathway
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...
Propagation of Waves
Consider a scenario where a wave propagates from a string of low linear mass density to a string of high linear mass density. In such a case, the reflected wave is out of phase with respect to the incident wave, however the...
Interference: Path Lengths
Two special sources may be considered when they are in phase. This can be easily achieved by feeding the two sources from the same source. An example would be synchronizing the two speakers by feeding them with the same source, such as the sound waves produced by a tuning fork. This setup ensures that the two sources have the same frequency and are...
Neuronal Communication

