Related Experiment Videos
Tone recognition of continuous Mandarin speech assisted with prosodic information
1Department of Communication Engineering, National Chiao Tung University, Hsinchu, Taiwan, Republic of China.
The Journal of the Acoustical Society of America
|November 1, 1994
Summary
A simple recurrent neural network (SRNN) improves Mandarin speech tone recognition by modeling prosody. This method enhances accuracy by using acoustic and linguistic features to predict tonal states, aiding speech technology development.
Area of Science:
- Computational Linguistics
- Speech Processing
- Artificial Intelligence
Background:
- Accurate tone recognition is crucial for understanding tonal languages like Mandarin.
- Traditional methods often struggle with the complex prosodic variations in continuous speech.
Purpose of the Study:
- To develop a novel approach for Mandarin speech tone recognition using a simple recurrent neural network (SRNN).
- To model the prosody of continuous Mandarin speech to enhance tone recognition accuracy.
Main Methods:
- Extracted acoustic and linguistic features from continuous Mandarin speech syllables.
- Employed a simple recurrent neural network (SRNN) to model prosodic states based on these features.
- Utilized SRNN hidden node outputs as additional features for a multilayer perception (MLP)-based tone recognition system.
Main Results:
- Achieved an improvement in speaker-dependent tone recognition rate from 91.38% to 93.10%.
- Demonstrated the SRNN's capability to learn and represent prosodic states effectively.
- Obtained a finite-state automata from SRNN outputs, offering insights into human prosody mechanisms.
Conclusions:
- The proposed SRNN-based prosody modeling significantly enhances Mandarin tone recognition.
- SRNNs offer a promising approach for capturing dynamic prosodic information in speech.
- Further analysis of SRNN models can reveal linguistic insights into prosody generation.