Related Experiment Video
Updated: Aug 5, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Exploring Prosodic Features Modelling for Secondary Emotions Needed for Empathetic Speech Synthesis.
Jesin James1, Balamurali B T2, Catherine Watson1
1Department of Electrical, Computer, and Software Engineering, The University of Auckland, Auckland 1010, New Zealand.
This study introduces a novel low-resource system for synthesizing secondary emotions in speech, crucial for empathetic communication. The approach effectively models subtle emotional nuances using machine learning, achieving over 65% accuracy in emotion recognition.
Area of Science:
- Speech synthesis
- Computational linguistics
- Affective computing
Background:
- Empathetic speech synthesis requires modeling secondary emotions, which are subtle and challenging to capture.
- Existing methods often rely on large datasets and deep learning, making them resource-intensive.
- Secondary emotions in speech synthesis remain under-explored.
Purpose of the Study:
- To develop a low-resource system for synthesizing secondary emotions in speech.
- To model subtle emotional prosody features using machine learning.
- To create an emotional text-to-speech (TTS) system capable of generating five distinct secondary emotions.
Main Methods:
- Handcrafted feature extraction and low-resource machine learning for prosody modeling.
- Quantitative model-based transformation for fundamental frequency contour shaping.
- Rule-based approaches for modeling speech rate and mean intensity.
- Development of an emotional TTS system synthesizing anxious, apologetic, confident, enthusiastic, and worried emotions.
Main Results:
- A proof-of-concept system for synthesizing secondary emotions in speech was successfully developed.
- The system utilizes a low-resource-intensive approach, avoiding the need for extensive emotion-specific databases.
- A perception test demonstrated that participants could identify synthesized emotions with a hit rate exceeding 65%.
Conclusions:
- The proposed low-resource approach is effective for synthesizing secondary emotions in speech.
- This method offers a viable alternative to resource-intensive deep learning models for emotional TTS.
- The system demonstrates potential for creating more nuanced and empathetic synthetic voices.
More Related Videos
Related Concept Videos
Empathy
Labeling Emotion
Emotional Expression
Universal Facial Expressions
Psychologist Paul Ekman identified seven basic...
Nonconscious Mimicry
Facial Feedback Hypothesis
Physiology of Emotion
Autonomic Nervous System
The autonomic nervous system (ANS) plays a critical role in emotional responses by regulating involuntary physiological functions. It consists of two main components: the sympathetic and parasympathetic systems. The sympathetic system...

