Related Experiment Video
Updated: Nov 6, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Learning emotions latent representation with CVAE for text-driven expressive audiovisual speech synthesis
Sara Dahmani1, Vincent Colotte1, Valérian Girard1
1Université de Lorraine, CNRS, Inria, LORIA, F-54000 Nancy, France.
Researchers developed new methods for expressive audiovisual Text-to-Speech synthesis (EAVTTS) using deep learning. They successfully generated nuanced and novel emotional speech by manipulating continuous emotion representations with a conditional variational auto-encoder (CVAE).
Area of Science:
- Artificial Intelligence
- Speech Synthesis
- Machine Learning
Background:
- Deep learning has advanced expressive audiovisual Text-to-Speech synthesis (EAVTTS).
- Controlling speech variability and generating realistic emotional speech remain challenges.
- Existing methods often lack flexibility in emotional representation.
Purpose of the Study:
- To synthesize emotional speech using novel neural architectures.
- To explore unsupervised learning for emotional speech modeling.
- To develop a continuous and flexible emotion representation for generating mixed and novel emotional speech styles.
Main Methods:
- Development and validation of an expressive audiovisual corpus.
- Analysis of a fully connected neural network for emotion-specific feature learning (phone duration, acoustic, visual modalities).
- Application of a conditional variational auto-encoder (CVAE) for unsupervised latent emotion representation and manipulation.
Main Results:
- Perceptual experiments validated the corpus's emotional content across acoustic, visual, and audiovisual stimuli.
- The fully connected network effectively learned emotion-specific characteristics.
- CVAE enabled generation of speech nuances and novel emotions with coherent articulation by manipulating latent vectors.
Conclusions:
- Unsupervised learning and continuous emotion representation are effective for EAVTTS.
- CVAE facilitates the generation of diverse and novel emotional speech styles.
- The proposed methods significantly enhance the control and realism of synthetic emotional speech.
More Related Videos
Related Concept Videos
Elaborative Rehearsals
The effectiveness of...
Emotional Expression
Universal Facial Expressions
Psychologist Paul Ekman identified seven basic...
Labeling Emotion
Language and Cognition
Facial Feedback Hypothesis
Chunking and Rehearsal in Sensory Memory

