Related Experiment Video
Updated: Aug 11, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.6K
Enhancing the Performance of Pathological Voice Quality Assessment System Through the Attention-Mechanism Based
Ji-Yan Han1, Ching-Ju Hsiao1, Wei-Zhong Zheng1
1National Yang Ming Chiao Tung University, Department of Biomedical Engineering, Taipei, Taiwan.
Journal of Voice : Official Journal of the Voice Foundation
|February 2, 2023
Summary
A new self-attention-based bidirectional long-short term memory (SA BiLSTM) system offers more accurate voice quality assessment than traditional methods. This deep learning approach improves pathological voice evaluation for better patient outcomes in voice therapy.
Area of Science:
- Computational Linguistics
- Artificial Intelligence
- Speech Pathology
Background:
- Current voice quality assessment relies on subjective auditory-perceptual evaluations by physicians, leading to variability in diagnosis.
- Subjectivity and time intervals in diagnosis can impact the accuracy of pathological voice assessment.
- An accurate, automated system is needed to enhance the reliability and efficiency of voice quality evaluations.
Purpose of the Study:
- To develop and evaluate a novel computerized system for pathological voice quality assessment.
- To improve the accuracy and consistency of voice disorder diagnosis using deep learning.
- To create a system that mimics professional doctors' evaluation of voice parameters.
Main Methods:
- A self-attention-based bidirectional long-short term memory (SA BiLSTM) deep learning model was developed.
- The model incorporated various pitches (low, normal, high) and vowels (/a/, /i/, /u/) to enhance learning.
- The system was trained to learn the complex, high-dimensional aspects of voice quality evaluation.
Main Results:
- The SA BiLSTM system demonstrated superior performance compared to baseline deep neural network and convolution neural network systems.
- The macro average F1 scores for grade, roughness, and breathiness were significantly higher with the proposed system (0.768, 0.820, 0.815) than baseline systems.
- The experimental results confirm the enhanced accuracy of the SA BiLSTM model in classifying voice quality parameters.
Conclusions:
- The proposed SA BiLSTM system, utilizing specific pitches and vowels, provides a more accurate method for voice quality assessment.
- This advanced system can significantly aid clinical voice evaluations, leading to improved patient benefits from voice therapy.
- The findings suggest a promising future for AI-driven tools in objective pathological voice analysis.

