Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Evaluation of Mixed Deep Neural Networks for Reverberant Speech Enhancement.

Biomimetics (Basel, Switzerland)·2019
Same author

Improving Post-Filtering of Artificial Speech Using Pre-Trained LSTM Neural Networks.

Biomimetics (Basel, Switzerland)·2019
See all related articles

Related Experiment Video

Updated: Nov 18, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.8K

Discriminative Multi-Stream Postfilters Based on Deep Learning for Enhancing Statistical Parametric Speech Synthesis.

Marvin Coto-Jiménez1

  • 1Electrical Engineering Department, University of Costa Rica, San José 11501-2060, Costa Rica.

Biomimetics (Basel, Switzerland)
|February 10, 2021
PubMed
Summary

This study introduces discriminative postfilters using long short-term memory (LSTM) deep neural networks to enhance Hidden Markov Model (HMM) synthesized speech quality. The new method improves intelligibility and naturalness, especially for low-resource languages.

Keywords:
deep learninglstmpostfilteringspeech synthesis

More Related Videos

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

668
Synthetic, Multi-Layer, Self-Oscillating Vocal Fold Model Fabrication
10:16

Synthetic, Multi-Layer, Self-Oscillating Vocal Fold Model Fabrication

Published on: December 2, 2011

14.3K

Related Experiment Videos

Last Updated: Nov 18, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.8K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

668
Synthetic, Multi-Layer, Self-Oscillating Vocal Fold Model Fabrication
10:16

Synthetic, Multi-Layer, Self-Oscillating Vocal Fold Model Fabrication

Published on: December 2, 2011

14.3K

Area of Science:

  • Speech synthesis
  • Artificial intelligence
  • Machine learning

Background:

  • Statistical parametric speech synthesis using Hidden Markov Models (HMM) offers high intelligibility and efficiency for low-resource languages.
  • However, HMM-based speech quality lags behind deep learning approaches.
  • Postfiltering is a proposed method to improve HMM speech quality.

Purpose of the Study:

  • To introduce a novel postfiltering technique for HMM-based synthesized voices.
  • To enhance speech quality by applying discriminative postfilters based on long short-term memory (LSTM) deep neural networks.
  • To specifically address distinct degradations in voiced and unvoiced sound segments.

Main Methods:

  • Development of discriminative postfilters utilizing multiple LSTM deep neural networks.
  • Modeling the mapping from synthesized to natural speech for voiced and unvoiced segments separately.
  • Evaluation using five distinct voices, Mel cepstral distance, and subjective listening tests.

Main Results:

  • Discriminative postfilters demonstrated significant advantages over standard HTS voices.
  • The proposed method outperformed non-discriminative postfilters in speech quality enhancement.
  • Objective and subjective evaluations confirmed the effectiveness of the approach.

Conclusions:

  • Discriminative postfiltering with LSTMs offers a promising avenue for improving HMM-based speech synthesis.
  • This technique effectively addresses quality degradation in synthesized speech, particularly for voiced/unvoiced segments.
  • The approach provides a viable solution for enhancing artificial voices, especially in low-resource scenarios.