Related Experiment Video
Updated: Jan 14, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Severity-Controllable Pathological Text-to-Speech Synthesis for Clinical Applications
Abstract:
The article presents a new pathological text-to-speech (TTS) synthesis system that has the ability to control speech severity using latent interpolations. Recognizing the difficulty of this task, our work uses a data augmentation technique to generate a single-speaker multi-severity training dataset required for training such a model. Furthermore, we show how x-vectors already contain information about the severity and leverage it as a conditioning variable for the synthesis. Finally, we propose modifications to the GradTTS architecture to enhance the duration modeling of pathological speech. We carry out objective and subjective evaluations to demonstrate that the proposed GradTTS system works well, and produces more natural, controllable, and stable pathological speech samples than the baseline TransformerTTS system.
More Related Videos
Related Concept Videos
Auditory Pathway
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...
Master Transcription Regulators
Master Transcription Regulators
Frequency-Domain Interpretation of PD Control
The proportional control gain, combined with the...

