Related Experiment Video
Updated: May 29, 2025

Author Spotlight: Advancements in the Fabrication of Synthetic Vocal Fold Models for Phonetic and Robotic Applications
Published on: January 5, 2024
A Deep-Learning Model for Multi-class Audio Classification of Vocal Fold Pathologies in Office Stroboscopy
Yeo E Kim1, Maria Dobko1, Haomiao Li2
1Sean Parker Institute for the Voice, Department of Otolaryngology-Head and Neck Surgery, Weill Cornell Medicine, New York, New York, U.S.A.
Deep learning models using voice data from stroboscopy videos showed moderate accuracy in distinguishing healthy vocal folds (VF) from unilateral paralysis (UVFP) and VF lesions, but struggled with multi-class classification and external validation.
Area of Science:
- Otolaryngology
- Artificial Intelligence in Medicine
- Speech Science
Background:
- Videolaryngostroboscopy is a key tool for visualizing vocal fold (VF) dynamics.
- Differentiating between healthy VFs, unilateral paralysis (UVFP), and VF lesions is crucial for diagnosis and treatment.
- Automated classification of VF conditions using voice analysis holds potential for improved diagnostic accuracy.
Purpose of the Study:
- To develop and validate deep-learning classifiers for distinguishing three VF states: healthy (HVF), UVFP, and VF lesions.
- To assess the performance of binary (HVF vs. pathological) and multi-class (HVF, UVFP, lesions) classification models.
- To evaluate model generalizability on an independent external dataset.
Main Methods:
- Voice data were extracted from stroboscopic videos of 105 UVFP, 63 VF lesion, and 41 HVF patients.
- Audio samples were converted to Mel-spectrograms and used to train ResNet18 deep-learning models.
- Models were trained for binary and multi-class classification and validated internally and externally.
Main Results:
- The binary classifier achieved higher performance on the test set (accuracy 83%, F1-score 0.90) than the multi-class classifier (accuracy 40%, F1-score 0.36).
- External validation showed reduced performance for both models, with the binary classifier achieving 63% accuracy and 0.48 F1-score.
- The multi-class classifier performed poorly on the external dataset (accuracy 35%, F1-score 0.25).
Conclusions:
- Deep learning models can differentiate healthy and pathological vocal fold conditions from stroboscopic video voice data with moderate accuracy.
- Multi-class classification and external validation significantly reduced model performance, indicating challenges in generalization.
- Voice data from stroboscopic recordings may have limited standalone diagnostic value for complex VF conditions; further research is warranted.
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Ultrasound II: Endoscopic Ultrasound and FibroScan
Endoscopic Ultrasound (EUS):
Classification of Epithelial Tissues: Stratified Epithelium
Endoscopic Studies I: Bronchoscopy and Thoracoscopy
Bronchoscopy
Description
Bronchoscopy is a procedure that involves direct visualization of the larynx, trachea, and bronchi for diagnostic and therapeutic purposes. A flexible fiber optic or rigid bronchoscope is used to carry out the procedure. The fiber-optic bronchoscope is more frequently used due...
Respiratory System Abnormal Finding II: Palpation and Auscultation
Palpation Findings
During a respiratory assessment, palpation can reveal several vital abnormalities:
Auditory Pathway
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...

