Related Experiment Video
Updated: May 9, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
A hybrid approach for binary and multi-class classification of voice disorders using a pre-trained model and ensemble
Mehtab Ur Rahman1,2, Cem Direkoglu3
1Department of Language and Communication, Radboud University, Houtlaan, Nijmegen, Gelderland, 6525, Netherlands. mehtab.rahman@ru.nl.
This study introduces a novel hybrid artificial intelligence (AI) approach for classifying voice disorders. The AI model achieves state-of-the-art accuracy in multi-class voice disorder classification, outperforming existing methods.
Area of Science:
- Artificial Intelligence
- Speech Processing
- Bioacoustics
Background:
- Accurate classification of voice disorders is crucial for diagnosis and treatment.
- Current AI methods face challenges in achieving high accuracy, particularly for multi-class classification.
- Existing feature extraction techniques may not capture all relevant vocal characteristics.
Purpose of the Study:
- To propose and evaluate a novel two-stage hybrid AI framework for enhanced voice disorder classification.
- To achieve state-of-the-art accuracies in both binary and multi-class voice disorder classification.
- To compare the proposed method against established techniques using a standard voice database.
Main Methods:
- A two-stage hybrid approach combining deep learning features with traditional classifiers.
- Stage 1: Feature embedding extraction from voice spectrograms using a pre-trained VGGish model.
- Stage 2: Classification using Support Vector Machine (SVM), Logistic Regression (LR), Multi-Layer Perceptron (MLP), and Ensemble Classifier (EC).
Main Results:
- The hybrid VGGish-SVM model achieved the highest accuracy in multi-class classification for male (77.81%) and combined (70.53%) speakers.
- For binary classification, VGGish-SVM and VGGish-EC showed top performance for male and female speakers, respectively.
- The proposed hybrid approach consistently outperformed MFCC, MFCC-glottal, wav2vec, and HuBERT-based methods.
Conclusions:
- The novel hybrid AI framework demonstrates significant potential for improving voice disorder classification accuracy.
- The VGGish feature embeddings combined with classifiers offer a robust approach for complex voice analysis.
- This work provides a foundation for developing advanced automated tools to aid clinical voice disorder assessment.
Related Concept Videos
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Classification of Systems-II
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...

