Related Experiment Video
Updated: Jan 18, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Auditory Feature Extraction Approach for Robust Pathological Voice Recognition
Youssef Zouhir1, Mohamed Zarka1, Lilia El Amraoui2
1Research Laboratory Smart Electricity & ICT, SE&ICT Lab, LR18ES44, National Engineering School of Carthage, University of Carthage, Tunis, Tunisia.
None:
While binary discrimination between healthy and pathological voices provides initial screening capability, clinical utility requires accurate multi-class classification to distinguish between different pathology types. This paper presents a novel Auditory Feature Extraction (AFE) approach, inspired by auditory perception mechanisms, for a robust multi-class pathological voice recognition (PVR) system. The proposed approach uses a Gammachirp FilterBank (GCFB) to accurately simulate the spectral behavior of the cochlea. This FilterBank is considered an efficient auditory filter model and contains 128 Gammachirp auditory filters with center frequencies equally distributed according to the Equivalent Rectangular Bandwidth-rate scale. The GCFB output undergoes decimation, cubic-root amplitude compression, and then Discrete Cosine Transform to finally generate the AFE coefficients. We evaluated the performance of the proposed AFE approach for PVR using Hidden Markov Model Toolkit on two widely recognized datasets: the Saarbruecken Voice Database (SVD) and the Massachusetts Eye and Ear Infirmary (MEEI) dataset. Our AFE outperforms state-of-the-art approaches as revealed by assessment results compared to Human Factor Cepstral Coefficients (HFCC), Frequency Domain Linear Prediction (FDLP), and Mel-Frequency Cepstral Coefficients (MFCC). AFE approach achieves 99.75% balanced accuracy in binary voice classification (95.6%, 94.8%, and 93.85% for HFCC, FDLP, and MFCC, respectively) and 94.38% balanced accuracy in multi-class pathologic voice classification (72.93%, 69.66%, and 60.03% for HFCC, FDLP, and MFCC, respectively) on the SVD dataset, while accomplishing 100% balanced accuracy on the MEEI database. These findings suggest the AFE approach provides a robust and highly discriminative feature set that could lead to an improved voice pathology classification, then a better clinical screening of voice disorders.
Related Concept Videos
Auditory Pathway
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...
Respiratory System Abnormal Finding II: Palpation and Auscultation
Palpation Findings
During a respiratory assessment, palpation can reveal several vital abnormalities:
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Auditory Perception
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Physical Assessment of the Respiratory Tract IV: Auscultation
Breath Sounds
Breath sounds are categorized into vesicular, bronchovesicular, and bronchial.

