Related Experiment Video
Updated: May 7, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Emotional speech acoustic model for Malay: iterative versus isolated unit training
Mumtaz Begum Mustafa1, Raja Noor Ainon
1Department of Software Engineering, Faculty of Computer Science and Information Technology, University of Malaya, 50603 Kuala Lumpur, Malaysia.
This study demonstrates building an emotional speech acoustic model for Malay with limited resources. It uses a neutral speech model and adaptation techniques for better emotional speech synthesis, even for under-resourced languages.
Area of Science:
- Speech synthesis
- Computational linguistics
- Acoustic modeling
Background:
- Emotional speech synthesis enhances user experience but is complex.
- State-of-the-art systems require extensive resources, posing challenges for under-resourced languages like Malay.
- Building emotional speech acoustic models necessitates segment-phonetic labels, often unavailable for many languages.
Purpose of the Study:
- To develop an emotional speech acoustic model for Malay using minimal resources.
- To investigate effective initialization and transformation techniques for low-resource emotional speech synthesis.
- To evaluate the naturalness and accuracy of the synthesized Malay emotional speech.
Main Methods:
- Utilized iterative training with the deterministic annealing expectation maximization algorithm and isolated unit training for model initialization.
- Employed a neutral speech acoustic model as a seed, transformed using model adaptation and context-dependent boundary refinement.
- Conducted objective evaluations for prosody error and subjective listening tests for naturalness.
Main Results:
- Successfully developed an emotional speech acoustic model for Malay with minimal data.
- Demonstrated the effectiveness of model adaptation and context-dependent boundary refinement techniques.
- Achieved acceptable naturalness in synthesized emotional speech, validated by objective and subjective evaluations.
Conclusions:
- It is feasible to build emotional speech acoustic models for under-resourced languages like Malay with limited resources.
- The proposed methods offer a viable approach for developing emotional speech synthesis systems in data-scarce environments.
- Further research can explore expanding these techniques to other under-resourced languages and emotional expressions.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
09:27Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018