Related Experiment Video
Updated: Jul 12, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
End-to-end deep learning classification of vocal pathology using stacked vowels
George S Liu1,2, Jordan M Hodges3, Jingzhi Yu4
1Department of Otolaryngology Head and Neck Surgery Stanford University School of Medicine, Stanford University Stanford California USA.
Analyzing multiple vowel recordings simultaneously with artificial intelligence (AI) improves voice disorder classification. The stacked vowel model shows promise for enhanced AI-driven screening of vocal pathology.
Area of Science:
- Computational linguistics
- Artificial intelligence in healthcare
- Speech pathology
Background:
- Artificial intelligence (AI) technology offers potential for classifying voice disorders using voice recordings as a screening tool.
- Previous models often rely on single vowel recordings, limiting prediction accuracy for vocal pathology.
Purpose of the Study:
- To develop and evaluate AI models that analyze multiple sustained vowel recordings simultaneously to enhance the prediction of vocal pathology.
- To compare the performance of a stacked vowel model against baseline and stacked pitch models for classifying voice disorders.
Main Methods:
- Voice samples from 687 healthy participants and 334 dysphonic patients (hyperfunctional dysphonia, laryngitis) were analyzed.
- Three 1-dimensional convolutional neural network models were trained: a baseline (single vowel), a stacked vowel model (three vowels simultaneously), and a stacked pitch model (one vowel across three pitches).
Main Results:
- The stacked vowel model achieved a higher F1 score (0.81) for multiclass classification compared to the baseline (0.77) and stacked pitch (0.78) models.
- The stacked vowel model demonstrated superior performance in classifying hyperfunctional dysphonia voice samples (F1 score 0.56) compared to other models.
Conclusions:
- Analyzing multiple sustained vowel recordings simultaneously significantly improves AI-driven screening and classification of vocal pathology.
- The stacked vowel model architecture presents a promising approach for enhancing the accuracy of AI-based voice disorder detection.
Related Concept Videos
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Larynx
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids,...
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Auditory Pathway
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...

