Related Experiment Videos
Speech synthesis by glottal excited linear prediction
1Department of Electrical Engineering, University of Florida, Gainesville 32611-2024.
The Journal of the Acoustical Society of America
|October 1, 1994
Summary
Quantizing glottal excitation waveforms does not degrade speech quality in the new glottal excitation linear predictive (GELP) synthesizer. This method offers high-quality speech synthesis with potential for voice analysis and modification.
Area of Science:
- Speech synthesis
- Digital signal processing
- Acoustic phonetics
Background:
- Linear predictive (LP) speech synthesis is a common technique.
- Modeling glottal excitation is crucial for realistic speech synthesis.
- Existing methods may not fully capture glottal waveform nuances.
Purpose of the Study:
- To demonstrate that quantizing the glottal excitation waveform does not significantly degrade synthesized speech quality.
- To introduce a novel Glottal Excitation Linear Predictive (GELP) synthesizer.
- To explore the potential of a codebook-based glottal excitation model.
Main Methods:
- Developed a 6th-order polynomial waveform model for glottal excitation.
- Designed and trained a 32-entry glottal excitation codebook for voiced sounds.
- Utilized a 256-entry stochastic codebook for unvoiced noise excitation.
- Implemented a GELP synthesizer based on pitch-excited LP and CELP coder principles.
Main Results:
- The GELP synthesizer resynthesizes speech with high quality.
- Quantization of the glottal excitation waveform showed minimal impact on speech quality.
- The excitation model waveform relates to the derivative of glottal flow and integral of the residue.
Conclusions:
- The GELP synthesizer offers a simple, high-quality speech synthesis procedure.
- The approach can reproduce all speech sounds using established analysis techniques.
- The glottal excitation codebook enables quantitative comparison of different voice types and potential voice conversion.