Related Experiment Video
Updated: Mar 12, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Improving zero-shot style transfer text-to-speech by disentangled fine-grained style modeling
Eray Eren1, Qingju Liu2, Abeer Alwan1
1Department of Electrical and Computer Engineering, University of California, Los Angeles, California 90095, USA.
This study introduces a novel zero-shot style-transfer speech synthesis method using a small GenerSpeech model. The approach enhances style adaptation for unseen speaking styles, outperforming existing methods.
Area of Science:
- Speech Synthesis
- Artificial Intelligence
- Machine Learning
Background:
- Zero-shot style-transfer speech synthesis aims to adapt to unseen speaking styles.
- Current state-of-the-art methods often require large models or extensive data for generalization.
- Achieving effective generalization with smaller models remains a significant challenge.
Purpose of the Study:
- To propose a novel zero-shot style-transfer speech synthesis method utilizing a compact GenerSpeech backbone.
- To enhance the adaptation capabilities of speech synthesis models to novel speaking styles with limited resources.
- To disentangle speaker, style, and content embeddings for improved synthesis quality.
Main Methods:
- A fine-grained style encoder was integrated with the GenerSpeech backbone.
- Mutual-information minimization loss was employed to disentangle speaker, global/fine-grained styles, and content embeddings.
- A maximum-mean-discrepancy-guided cycle consistency loss was introduced to further disentangle style from speaker and increase style embedding diversity.
Main Results:
- The proposed method demonstrated superior performance compared to established zero-shot style-transfer baselines (GenerSpeech, YourTTS, VALL-E-X).
- Achieved a relative average style preference improvement of 31%.
- Obtained a mean opinion score of 3.64 for prosody similarity on the VCTK dataset.
Conclusions:
- The developed method effectively achieves zero-shot style-transfer speech synthesis with a small model.
- The proposed disentanglement losses significantly improve style adaptation and synthesis quality.
- This work advances the capability of small models in generalizing to diverse speaking styles.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Impression Management Techniques IV: Altercasting
Stereotype Content Model
Modeling and Similitude
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
