Related Experiment Video
Updated: Jan 24, 2026

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
Improving Post-Filtering of Artificial Speech Using Pre-Trained LSTM Neural Networks
1Escuela de Ingeniería Eléctrica, Universidad de Costa Rica, San José 11501-2060, Costa Rica. marvin.coto@ucr.ac.cr.
This study introduces a novel pre-training method for Long Short-term Memory (LSTM) networks to improve statistical parametric speech synthesis quality. The auto-associative pre-training enhances Mel-Frequency Cepstral parameters more effectively than random initialization.
Area of Science:
- Speech Synthesis
- Deep Learning
- Signal Processing
Background:
- Deep learning post-filters aim to improve statistical parametric speech synthesis by mapping synthetic to natural speech.
- Long Short-term Memory (LSTM) networks are effective but have room for improvement in quality and efficiency.
- Current methods often use random initialization for LSTM networks in speech synthesis.
Purpose of the Study:
- To introduce a new pre-training approach for LSTM networks to enhance synthesized speech quality, particularly spectral characteristics.
- To improve the efficiency of speech synthesis enhancement processes.
- To evaluate the effectiveness of auto-associative pre-training for LSTM-based speech synthesis.
Main Methods:
- Implemented an auto-associative pre-training method for a single LSTM network.
- Utilized the pre-trained LSTM as an initialization strategy for post-filters in speech synthesis.
- Focused on enhancing Mel-Frequency Cepstral (MFC) parameters of synthetic speech.
Main Results:
- The proposed auto-associative pre-training initialization demonstrated advantages in enhancing MFC parameters.
- The pre-training approach achieved superior results in improving the statistical parametric speech spectrum compared to random initialization.
- Effectiveness was observed in most tested scenarios.
Conclusions:
- Auto-associative pre-training is a viable and effective method for initializing LSTMs in speech synthesis.
- This approach offers a more efficient way to enhance spectral quality in synthesized speech.
- The findings suggest a significant improvement over traditional random initialization techniques.
Related Concept Videos
Cardiomyopathy VII: Pre and Post Operative Nursing Management
Passive Filters
Low-Pass Filters
Low-pass filters are designed to transmit signals with frequencies lower than the cutoff frequency, ωc, and attenuate those above it. The cutoff...
Active Filters
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
pre-mRNA Processing
Once about 20-40 ribonucleotides have been joined together by RNA polymerase, a group of enzymes adds a “cap” to the 5’ end of the growing transcript. In this process, a 5’ phosphate is replaced by modified guanosine that has a methyl group attached to it (7-Methyl...
Network Covalent Solids
To break or to melt a covalent network solid, covalent bonds must be broken. Because covalent bonds are relatively strong, covalent network solids are typically...

