Related Experiment Video
Updated: Jul 1, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Voice Synthesis Improvement by Machine Learning of Natural Prosody.
Joseph Kane1,2, Michael N Johnstone1,2, Patryk Szewczyk1,2
1Cyber Security Cooperative Research Centre, Edith Cowan University, 270 Joondalup Drive, Joondalup, WA 6027, Australia.
Researchers developed a novel machine learning approach using LSTM neural networks to add realistic prosody to computer-generated speech. This enhances human-computer interaction and artificial voices, making synthesized speech more natural.
Area of Science:
- Computer Science
- Artificial Intelligence
- Speech Synthesis
Background:
- Current human-computer interfaces (HCI) strive for seamless interaction.
- Computer-generated speech often lacks natural prosody (intonation and rhythm), sounding non-human.
- Improving synthesized voice realism is crucial for applications like electronic readers and assistive speaking devices.
Purpose of the Study:
- To enhance computer-generated text-to-speech (TTS) algorithms by incorporating human-like melodic and prosodic elements.
- To increase the realism of synthesized voices.
- To improve electronic reading applications and artificial voices for individuals needing speech assistance.
Main Methods:
- Exploration of a novel approach using machine learning, specifically a Long Short-Term Memory (LSTM) neural network.
- Implementation of an LSTM-based encoder to add paralinguistic elements to recorded or generated voices.
- Deployment of a prototype modular platform for digital speech improvement through laboratory experiments.
Main Results:
- The LSTM-based encoder demonstrated the ability to produce realistic speech.
- Laboratory experiments provided encouraging results for the developed speech improvement algorithms.
- Optimization of algorithm combinations and performance in edge cases was explored.
Conclusions:
- The novel LSTM-based approach shows significant promise for generating more natural and realistic computer-generated speech.
- Further research will focus on algorithm optimization and comparative performance analysis against existing methods.
- The developed technology has potential applications in improving audio codecs, restoring old recordings, and enhancing HCI.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
06:24Author Spotlight: Advancements in the Fabrication of Synthetic Vocal Fold Models for Phonetic and Robotic Applications
Published on: January 5, 2024
Related Concept Videos
Improving Translational Accuracy
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Non-equilibrium in the Cell
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...