Age estimation based on children's voice: a fuzzy-based decision fusion strategy
Seyed Mostafa Mirhassani1, Alireza Zourmand1, Hua-Nong Ting1
1Biomedical Engineering Department, Faculty of Engineering, University of Malaya, Lembah Pantai, 50603 Kuala Lumpur, Malaysia.
Thescientificworldjournal
|July 10, 2014
Summary
This study introduces a novel speech analysis method for automatic speaker age estimation. By dividing speech data by vowel class and using fuzzy fusion, age estimation accuracy improved by over 53%.
Area of Science:
- Speech Analysis
- Machine Learning
- Signal Processing
Background:
- Automatic speaker age estimation is a complex challenge in speech analysis.
- Existing methods face difficulties in accurately determining age from vocal characteristics.
Purpose of the Study:
- To present a novel, accurate approach for automatic speaker age estimation.
- To leverage a "divide and conquer" strategy using vowel classes and fuzzy data fusion.
Main Methods:
- Speech data were segmented into six groups based on vowel classes.
- Mel-frequency cepstral coefficients (MFCCs) were extracted for each group.
- Single-layer feed-forward neural networks (SLFNs) and self-adaptive extreme learning machines (ELMs) were used for classification, followed by fuzzy data fusion.
Main Results:
- The proposed method demonstrated significant improvements in age estimation accuracy, reaching up to 53.33%.
- Fuzzy fusion effectively aggregated complementary age-related information from different vowel classes.
- The approach outperformed several state-of-the-art age estimation techniques.
Conclusions:
- The "divide and conquer" strategy combined with fuzzy fusion enhances speaker age estimation accuracy.
- Vowel-class-based speech segmentation and fusion of classifier outputs are effective for age estimation.
- This method offers a robust solution for automatic speaker age determination across various age groups, including children.
Related Concept Videos
Perceiving Loudness, Pitch, and Location
1.3K
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
1.3K
Determination of Expected Frequency
1.7K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
1.7K


