Related Experiment Video
Updated: Sep 18, 2025

Minimally Invasive Murine Laryngoscopy for Close-Up Imaging of Laryngeal Motion During Breathing and Swallowing
Published on: December 1, 2023
Machine learning based assessment of hoarseness severity: a multi-sensor approach centered on high-speed
Tobias Schraut1, Anne Schützenberger1, Tomás Arias-Vergara2
1Division of Phoniatrics and Pediatric Audiology at the Department of Otorhinolaryngology, Head and Neck Surgery, University Hospital Erlangen, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany.
This study developed a machine learning model using high-speed videoendoscopy (HSV) and acoustic recordings to objectively assess hoarseness severity. Combining all data modalities provided the most accurate assessment of vocal function.
Area of Science:
- Otolaryngology
- Speech-Language Pathology
- Biomedical Engineering
Background:
- Functional voice disorders lack objective assessment due to reliance on subjective evaluations with high inter-rater variability.
- High-speed videoendoscopy (HSV) offers objective insights into vocal fold dynamics, complementing acoustic and clinical data.
- Synchronous recording of HSV and acoustic data enables a unified, objective assessment of voice function.
Purpose of the Study:
- To investigate a machine learning approach for objective hoarseness severity assessment.
- To evaluate the diagnostic value of combining high-speed videoendoscopy (HSV), acoustic, and clinical data.
- To compare the performance of individual and combined data modalities in assessing vocal function.
Main Methods:
- Analysis of 457 HSV recordings, 634 synchronized acoustic recordings, and clinical data from 923 visits.
- Development of a videoendoscopic model using machine learning for hoarseness classification and severity rating.
- Comparison of videoendoscopic, acoustic, and clinical models, including ensemble models combining modalities.
Main Results:
- Individual models showed moderate correlations with subjective hoarseness ratings (videoendoscopic: 0.464, acoustic: 0.512, clinical: 0.638).
- An ensemble model integrating HSV and acoustic data improved correlation to 0.603.
- A multimodal ensemble model incorporating all data types achieved the highest correlation (0.752) with subjective ratings.
Conclusions:
- Synchronous HSV and acoustic recordings offer a more comprehensive and objective assessment of vocal function.
- Multimodal data integration significantly enhances diagnostic accuracy for hoarseness severity.
- Future work should focus on flexible endoscopy and refined parameter extraction for improved robustness.
Related Concept Videos
Endoscopic Studies I: Bronchoscopy and Thoracoscopy
Bronchoscopy
Description
Bronchoscopy is a procedure that involves direct visualization of the larynx, trachea, and bronchi for diagnostic and therapeutic purposes. A flexible fiber optic or rigid bronchoscope is used to carry out the procedure. The fiber-optic bronchoscope is more frequently used due...
Endoscopic Procedures III: Video Capsule Endoscopy
Assessment of Airway, Skin Color, and Use of Accessory Muscles
Introduction
The initial evaluation of a patient's respiratory system...
Assessment of Ventilation II: Respiratory Depth and Rhythm
Respiratory depth measures the volume of air inhaled or exhaled during a breath. It can vary from shallow to deep and typically remains consistent when a person is at rest or asleep. Occasionally, individuals will automatically inhale deeply, known as sighing, which inflates the lungs with more air than normal breathing.
To assess respiratory depth, observe the degree of chest excursion or movement:

