Related Experiment Video
Updated: Sep 18, 2025

Minimally Invasive Murine Laryngoscopy for Close-Up Imaging of Laryngeal Motion During Breathing and Swallowing
Published on: December 1, 2023
Machine learning based assessment of hoarseness severity: a multi-sensor approach centered on high-speed
Tobias Schraut1, Anne Schützenberger1, Tomás Arias-Vergara2
1Division of Phoniatrics and Pediatric Audiology at the Department of Otorhinolaryngology, Head and Neck Surgery, University Hospital Erlangen, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany.
Introduction:
Functional voice disorders are characterized by impaired voice production without primary organic changes, posing challenges for standardized assessment. Current diagnostic methods rely heavily on subjective evaluation, suffering from inter-rater variability. High-speed videoendoscopy (HSV) offers an objective alternative by capturing true intra-cycle vocal fold behavior. Integrating time-synchronized acoustic and HSV recordings could allow for an objective visual and acoustic assessment of vocal function based on a single HSV examination. This study investigates a machine learning-based approach for hoarseness severity assessment using synchronous HSV and acoustic recordings, alongside conventional voice examinations.
Methods:
Three databases comprising 457 HSV recordings of the sustained vowel /i/, 634 HSV-synchronized acoustic recordings, and clinical parameters from 923 visits were analyzed. Subjects were classified into two hoarseness groups based on auditory-perceptual ratings, with predicted scores serving as continuous hoarseness severity ratings. A videoendoscopic model was developed by selecting a suitable classification algorithm and a minimal-optimal subset of glottal parameters. This model was compared against an acoustic model based on HSV-synchronized recordings and a clinical model based on parameters from other examinations. Two ensemble models were constructed by combining the HSV-based models and all models, respectively. Model performance was evaluated on a shared test set based on classification accuracy, correlation with subjective ratings, and correlation between predicted and observed changes in hoarseness severity.
Results:
The videoendoscopic, acoustic, and clinical model achieved correlations of 0.464, 0.512, and 0.638 with subjective hoarseness ratings. Integrating glottal and acoustic parameters into the HSV-based ensemble model improved correlation to 0.603, confirming the complementary nature of time-synchronized HSV and acoustic recordings. The ensemble model incorporating all modalities achieved the highest correlation of 0.752, underscoring the diagnostic value of multimodal objective assessments.
Discussion:
This study highlights the potential of synchronous HSV and acoustic recordings for objective hoarseness severity assessment, offering a more comprehensive evaluation of vocal function. While practical challenges remain, the integration of these modalities led to notable improvements, supporting their complementary value in enhancing diagnostic accuracy. Future advancements could include flexible nasal endoscopy to enable more natural phonation and refinement of glottal parameter extraction to improve model robustness under variable recording conditions.
Related Concept Videos
Endoscopic Studies I: Bronchoscopy and Thoracoscopy
Bronchoscopy
Description
Bronchoscopy is a procedure that involves direct visualization of the larynx, trachea, and bronchi for diagnostic and therapeutic purposes. A flexible fiber optic or rigid bronchoscope is used to carry out the procedure. The fiber-optic bronchoscope is more frequently used due...
Endoscopic Procedures III: Video Capsule Endoscopy
Assessment of Airway, Skin Color, and Use of Accessory Muscles
Introduction
The initial evaluation of a patient's respiratory system...
Assessment of Ventilation II: Respiratory Depth and Rhythm
Respiratory depth measures the volume of air inhaled or exhaled during a breath. It can vary from shallow to deep and typically remains consistent when a person is at rest or asleep. Occasionally, individuals will automatically inhale deeply, known as sighing, which inflates the lungs with more air than normal breathing.
To assess respiratory depth, observe the degree of chest excursion or movement:

