Related Experiment Video
Updated: May 24, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Robust Sequence-to-sequence Voice Conversion for Electrolaryngeal Speech Enhancement in Noisy and Reverberant
Abstract:
Electrolaryngeal (EL) speech, an artificial speech produced by an electrolarynx for laryngectomees, lacks essential phonetic features, and differs in temporal structure from normal speech, resulting in poor naturalness and intelligibility. To address this deficiency, sequence-to-sequence (seq2seq) voice conversion (VC) models have been applied in converting EL speech to normal speech (EL2SP), showing some promising performances. However, previous studies mostly focus on converting clean EL speech, thereby restricting the further applicability in real-world scenarios, especially when the EL speech is inevitably interfered with background noise and reverberation. In light of this, we suggest novel training techniques based on seq2seq VC to enhance the robustness of real-world EL2SP. We first pretrain a normal-to-normal seq2seq VC model based on a text-to-speech model. Then, a two-stage fine-tuning is conducted by effectively using pseudo noisy and reverberant EL speech data artificially generated from only a small amount of original clean data available. Several design options are investigated to figure out the effectiveness of our method. The significant improvements presented in experimental results indicate that our method can non-trivially handle both clean and noisy-reverberant EL speech, enhancing the robustness of EL2SP in real-world scenarios.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
06:24Author Spotlight: Advancements in the Fabrication of Synthetic Vocal Fold Models for Phonetic and Robotic Applications
Published on: January 5, 2024
Related Concept Videos
Design Example: Vintage Mixing Console
The specifications for the pre-amplifier were clear. It needed to amplify the audio signal by a factor of 10, have an input impedance above 10...
Double Resonance Techniques: Overview
Spin decoupling is usually achieved by...
Lossy Lines and Overvoltages
Attenuation
When constant series resistance and shunt conductance are present, voltage and current equations are modified. The propagation constant indicates that voltage and current waves consist of both forward and backward traveling components. These waves attenuate as they propagate, with the attenuation factor related to the resistance and conductance. In a...
Design Example
Reconstruction of Signal using Interpolation