Related Experiment Videos
Zero-shot emotional speech synthesis based on feature decoupling and adaptive loss-threshold reweighting
Yan Zhu1, Yu Wang1, Yijin Zhou1
1School of Design and Art, Shanghai Dianji University, Shanghai, China.
Abstract:
Deep learning-based zero-shot speech synthesis has achieved substantial progress in speaker generalization, but stable modeling remains challenging in fine-grained emotional scenarios. Existing systems often process textual and emotional conditions through shared or closely coupled pathways, which may introduce interference between semantic content and emotional expression. In addition, uniformly averaging per-sample Conditional Flow Matching (CFM) losses may provide insufficient optimization emphasis to high-loss emotional samples. This study proposes a CFM-based zero-shot emotional speech synthesis method. Emotion-Text Decoupling Attention (ETDA) processes semantic and emotional conditions through parallel cross-attention streams and combines them through adaptive gated fusion, allowing the two conditions to retain their respective information before fusion. A 16-class fine-grained emotion space is constructed through classifier filtering and K-Means clustering based on pitch and energy features, and the resulting labels are mapped to continuous representations using a trainable lookup embedding. During training, Adaptive Loss-Threshold Reweighting estimates a threshold from mini-batch loss statistics and assigns larger weights to samples whose individual CFM losses exceed that threshold. Under speaker-disjoint evaluation on ESD, the proposed method achieves a word error rate of 3.62±0.06%, a mel-cepstral distortion of 4.45±0.05 dB, and an emotional expressiveness mean opinion score of 4.64±0.04. Controlled text-length expansion and analyses on a fixed subset of high-loss emotional samples further indicate that the proposed method maintains linguistic content and emotional expression more consistently under the evaluated conditions. These results suggest that the proposed method can, to some extent, improve content preservation, emotional expression, and acoustic reconstruction in fine-grained zero-shot emotional speech synthesis.
Related Concept Videos
Reducing Line Loss
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss in...
Downsampling
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
Line Loss
Line loss impacts power delivery efficiency in a balanced three-phase circuit. The symmetry in such a circuit simplifies the...
Lossy Lines and Overvoltages
Attenuation
When constant series resistance and shunt conductance are present, voltage and current equations are modified. The propagation constant indicates that voltage and current waves consist of both forward and backward traveling components. These waves attenuate as they propagate, with the attenuation factor related to the resistance and conductance. In a...
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear.
Reconstruction of Signal using Interpolation