Formant frequency estimation of high-pitched vowels using weighted linear prediction
Paavo Alku1, Jouni Pohjalainen, Martti Vainio
1Department of Signal Processing and Acoustics, Aalto University, P.O. Box 13000, FI-00076 Aalto, Finland. paavo.alku@aalto.fi
The Journal of the Acoustical Society of America
|August 10, 2013
Summary
This study introduces Weighted Linear Prediction with Attenuated Main Excitation (WLP-AME) to improve formant estimation for high-pitched voices. WLP-AME significantly reduces formant error compared to traditional methods, enhancing vocal tract analysis.
Area of Science:
- Speech processing
- Acoustic phonetics
- Signal processing
Background:
- All-pole modeling is a standard formant estimation technique.
- Performance degrades for high-pitched voices, impacting accuracy.
- Existing methods struggle with fundamental frequency variations.
Purpose of the Study:
- To compare existing fundamental frequency-robust all-pole modeling methods.
- To introduce and evaluate a novel technique, Weighted Linear Prediction with Attenuated Main Excitation (WLP-AME).
- To improve formant estimation accuracy for high-pitched voices.
Main Methods:
- Comparison of five established fundamental frequency-robust all-pole modeling techniques.
- Introduction of WLP-AME, utilizing temporally weighted linear prediction (LP).
- WLP-AME down-weights vocal tract excitation during filter coefficient optimization.
Main Results:
- WLP-AME demonstrated improved formant frequency accuracy for high-pitched synthetic vowels.
- Relative error for the first formant of vowel [a] reduced from 11% to 3% using WLP-AME.
- Natural vowel experiments showed more regular formant tracking with WLP-AME across pitch variations.
Conclusions:
- WLP-AME offers less biased formant estimates by emphasizing vocal tract characteristics.
- The method enhances the robustness of all-pole modeling for high-pitched speech.
- WLP-AME presents a significant advancement in accurate formant frequency detection.
Related Concept Videos
Linear Approximation in Frequency Domain
Linear systems are characterized by two main properties: superposition and homogeneity. Superposition allows the response to multiple inputs to be the sum of the responses to each individual input. Homogeneity ensures that scaling an input by a scalar results in the response being scaled by the same scalar.
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear.
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear.
Determination of Expected Frequency
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
Expected Frequencies in Goodness-of-Fit Tests
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
Linear Approximation in Time Domain
Nonlinear systems often require sophisticated approaches for accurate modeling and analysis, with state-space representation being particularly effective. This method is especially useful for systems where variables and parameters vary with time or operating conditions, such as in a simple pendulum or a translational mechanical system with nonlinear springs.
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length, the...
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length, the...
Perceiving Loudness, Pitch, and Location
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by identifying...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by identifying...
Mean From a Frequency Distribution
Sometimes, data gathered from an experiment on a large sample or population are organized into concise tables. In such cases, the frequency of the quantitative data set is plotted in the form of a table. Or else, the data values are grouped into the quantity’s intervals, which form classes, and their respective frequencies are known. That is, the data values are distributed over different categories or classes. This is known as frequency distribution.
When such a data set is encountered, the...
When such a data set is encountered, the...


