Related Experiment Video
Updated: Aug 19, 2025

04:04
Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
289
Automatic GRBAS Scoring of Pathological Voices using Deep Learning and a Small Set of Labeled Voice Data
Shunsuke Hidaka1, Yogaku Lee2, Moe Nakanishi1
1Graduate School of Design, Kyushu University, Fukuoka, Japan.
Journal of Voice : Official Journal of the Voice Foundation
|November 27, 2022
Summary
Deep learning models can accurately estimate pathological voice quality using the GRBAS scale, even with limited data. This approach matches expert reliability, overcoming data scarcity challenges in voice analysis.
Area of Science:
- Speech and Hearing Sciences
- Artificial Intelligence in Medicine
- Biomedical Signal Processing
Background:
- The grade-roughness-breathiness-asthenia-strain (GRBAS) scale is the standard for pathological voice quality assessment.
- Subjectivity and rater variability limit the reproducibility of traditional GRBAS evaluations.
- Existing deep learning methods for automatic GRBAS scoring require extensive labeled voice datasets.
Purpose of the Study:
- To investigate the feasibility of automatic GRBAS score estimation using deep learning with limited labeled voice data.
- To develop a neural network model capable of predicting GRBAS scores from voice waveforms.
- To evaluate the effectiveness of data augmentation and time-frequency representations for improving model performance.
Main Methods:
- A dataset of 300 pathological sustained /a/ vowel samples was curated and annotated by eight expert raters.
- A neural network model was designed to predict GRBAS score probability distributions from onset-to-offset waveforms.
- Data augmentation techniques (speed perturbation, crop, frequency masking) and time-frequency representations (power, instantaneous frequency, group delay) were explored.
Main Results:
- The proposed model achieved automatic scoring performance comparable to expert inter-rater reliability for all GRBAS items.
- Performance closely matched expert intra-rater reliability for items G, B, A, and S.
- Random speed perturbation proved to be the most effective data augmentation method; power was optimal for most items, with group delay enhancing performance for item B.
Conclusions:
- Deep learning models can achieve expert-level performance in automatic GRBAS scoring with limited data.
- The proposed methods effectively mitigate challenges associated with insufficient labeled voice data.
- This research contributes to advancing automatic voice disorder detection and related clinical applications.

