Related Experiment Videos
Predicting fundamental frequency from mel-frequency cepstral coefficients to enable speech reconstruction
1School of Computing Sciences, University of East Anglia, Norwich, NR4 7TJ, United Kingdom.
The Journal of the Acoustical Society of America
|September 15, 2005
Summary
This study reconstructs speech using only mel-frequency cepstral coefficients (MFCCs). It accurately predicts voicing and fundamental frequency from MFCCs, achieving speech quality comparable to methods using reference data.
Area of Science:
- Speech processing
- Acoustic signal analysis
- Machine learning for audio
Background:
- Distributed speech recognition (DSR) systems often rely on mel-frequency cepstral coefficients (MFCCs).
- Traditional speech reconstruction methods require MFCCs along with fundamental frequency and voicing information.
- Reconstructing speech solely from MFCCs presents a significant challenge.
Purpose of the Study:
- To develop a method for reconstructing acoustic speech signals exclusively from MFCCs.
- To predict voicing classification and fundamental frequency directly from MFCCs.
- To evaluate the effectiveness of the proposed speech reconstruction method.
Main Methods:
- Proposed a novel speech reconstruction technique utilizing only MFCCs.
- Employed two maximum a posteriori (MAP) methods for predicting voicing and fundamental frequency from MFCCs.
- Utilized a Gaussian mixture model (GMM) for joint density modeling of MFCCs and fundamental frequency.
- Implemented a Hidden Markov Model (HMM)-based approach with state-dependent GMMs for localized density modeling.
Main Results:
- Achieved accurate voicing classification and fundamental frequency prediction compared to reference measurements.
- Speaker-independent experiments on male and female speech demonstrated the method's efficacy.
- Speech reconstruction using predicted parameters yielded speech quality comparable to using reference parameters.
Conclusions:
- Accurate speech reconstruction is feasible using only MFCCs by predicting essential acoustic features.
- The proposed MAP-based methods effectively estimate voicing and fundamental frequency from MFCCs.
- This approach simplifies DSR systems by eliminating the need for separate voicing and fundamental frequency transmission.