Related Experiment Video
Updated: Mar 12, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Improving zero-shot style transfer text-to-speech by disentangled fine-grained style modeling
Eray Eren1, Qingju Liu2, Abeer Alwan1
1Department of Electrical and Computer Engineering, University of California, Los Angeles, California 90095, USA.
Abstract:
Recent zero-shot style-transfer speech synthesis methods have shown promising results and addressed adaptation to unseen speaking styles. While most state-of-the-art methods generalize to new speakers and styles using large models or corpora, achieving similar generalization with a smaller model remains an open challenge. We propose a zero-shot method that uses the small GenerSpeech backbone plus a fine-grained style encoder. To disentangle speakers, global/fine-grained styles, and content embeddings, we introduce a mutual-information minimization loss. To further disentangle style from speaker and boost style embedding diversity, we introduce a maximum-mean-discrepancy-guided cycle consistency loss. Experimental results show the proposed method outperforms baseline zero-shot style-transfer methods (GenerSpeech, YourTTS, VALL-E-X) with a relative average style preference improvement of 31% and a 3.64 prosody prosody similarity mean opinion score on VCTK.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Impression Management Techniques IV: Altercasting
Stereotype Content Model
Modeling and Similitude
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
