Related Experiment Video
Updated: Jan 1, 2026

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
968
Evaluation of Mixed Deep Neural Networks for Reverberant Speech Enhancement
Michelle Gutiérrez-Muñoz1, Astryd González-Salazar1, Marvin Coto-Jiménez1
1Escuela de Ingeniería Eléctrica, Universidad de Costa Rica, San José 11501-2060, Costa Rica.
Biomimetics (Basel, Switzerland)
|December 22, 2019
Summary
Hybrid neural networks offer efficient speech signal enhancement by reducing training time by 30% compared to pure long short-term memory (LSTM) networks, maintaining signal quality for voice recognition systems.
Area of Science:
- Artificial Intelligence
- Signal Processing
- Machine Learning
Background:
- Degraded speech signals in real-world environments pose challenges for voice recognition and analysis.
- Reverberation, caused by sound reflections, significantly degrades speech signal quality.
- Deep learning, particularly long short-term memory (LSTM) networks, shows promise for speech enhancement but suffers from high computational training costs.
Purpose of the Study:
- To evaluate hybrid neural network models for learning diverse reverberation conditions without prior information.
- To assess the effectiveness of combining LSTM and perceptron layers for speech signal enhancement.
- To compare the performance of hybrid models against pure LSTM networks in terms of quality, training time, and efficiency.
Main Methods:
- Trained and compared 120 artificial neural networks of eight different types, focusing on hybrid LSTM-perceptron architectures.
- Evaluated network performance using signal spectrum quality measurements and statistical validation.
- Measured and compared the training time of various network configurations.
Main Results:
- Certain hybrid LSTM-perceptron combinations achieved comparable results to pure LSTM networks with a fixed number of layers.
- Hybrid models demonstrated a significant reduction in training time, approximately 30%, compared to traditional LSTM approaches.
- The proposed hybrid networks offer efficiency gains without a substantial decrease in speech signal enhancement quality.
Conclusions:
- Hybrid neural networks are a viable and effective solution for speech signal enhancement in reverberant conditions.
- These models offer a practical alternative to pure LSTM networks, mitigating high computational costs and long training durations.
- The findings support the use of hybrid architectures for improving the robustness and efficiency of voice processing systems.

