Related Experiment Video
Updated: Jan 14, 2026

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
2.0K
Severity-Controllable Pathological Text-to-Speech Synthesis for Clinical Applications.
Summary
This study introduces a novel pathological text-to-speech (TTS) system for controllable speech synthesis. The system enhances pathological speech naturalness and stability using advanced data augmentation and architecture modifications.
Area of Science:
- Speech synthesis
- Pathological speech
- Machine learning
Background:
- Pathological speech synthesis presents challenges in controlling speech severity.
- Existing text-to-speech (TTS) systems struggle with generating diverse pathological speech characteristics.
- Generating multi-severity datasets for single-speaker pathological speech is difficult.
Purpose of the Study:
- To develop a novel pathological text-to-speech (TTS) synthesis system capable of controlling speech severity.
- To improve the naturalness, controllability, and stability of synthesized pathological speech.
- To leverage existing speaker embeddings (x-vectors) for severity conditioning.
Main Methods:
- Utilized data augmentation to create a single-speaker multi-severity training dataset.
- Employed x-vectors as a conditioning variable for speech severity control.
- Modified the GradTTS architecture to improve duration modeling for pathological speech.
Main Results:
- The proposed GradTTS system demonstrated superior performance compared to the baseline TransformerTTS.
- Objective and subjective evaluations confirmed the system's effectiveness.
- The system successfully generated more natural, controllable, and stable pathological speech samples.
Conclusions:
- The developed GradTTS system offers effective control over pathological speech severity.
- The integration of x-vectors and architectural modifications enhances pathological TTS capabilities.
- This work advances the field of controllable and natural-sounding pathological speech synthesis.
More Related Videos
Related Concept Videos
Auditory Pathway
7.1K
Auditory pathways constitute the complex neural circuits responsible for transmitting and interpreting auditory information from the peripheral auditory system to the brain. Sound waves are initially captured by the outer ear, funneled through the ear canal, and reach the tympanic membrane (eardrum). These vibrations are transmitted via the middle ear's ossicles to the inner ear's cochlea.
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...
7.1K
Master Transcription Regulators
2.7K
2.7K
Master Transcription Regulators
7.7K
Master transcription regulators are regulatory proteins that are predominantly responsible for regulating the expression of multiple genes. Often these genes work in concert to drive a complex process. Activation of a master transcription regulator can lead to a cascade of transcriptional activation necessary for that outcome. These regulators can directly bind to the regulatory sequences of the various genes involved, or they can indirectly regulate transcription by binding to regulatory...
7.7K
Frequency-Domain Interpretation of PD Control
346
Proportional-Derivative (PD) controllers are widely used in fan control systems to improve stability and performance. A fan control system can be effectively represented using a Bode plot to illustrate the impact of a PD controller through its transfer function. The Bode plot visually conveys how PD control modifies the fan's response across various frequencies, providing a frequency domain interpretation of the controller's behavior.
The proportional control gain, combined with the...
The proportional control gain, combined with the...
346

