Related Experiment Video
Updated: Jul 12, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Robust Vocal Quality Feature Embeddings for Dysphonic Voice Detection
Jianwei Zhang1, Julie Liss2, Suren Jayasuriya3
1School of Electrical, Computer and Energy Engineering, Arizona State University, Tempe, AZ 85281, USA.
This study introduces a novel deep learning framework for accurate automatic detection of voice impairments (dysphonia). The method generates robust acoustic embeddings, improving performance across different datasets and conditions.
Area of Science:
- Speech processing
- Biomedical engineering
- Machine learning
Background:
- Approximately 1.2% of the global population experiences impaired voice production, necessitating reliable automated assessment tools.
- Current automated voice analysis methods often lack generalizability across different datasets and applications.
- There is a significant need for robust and accurate methods for dysphonic voice detection.
Purpose of the Study:
- To develop a deep learning framework for generating acoustic feature embeddings that are sensitive to vocal quality.
- To enhance the robustness of voice analysis models across diverse corpora and conditions.
- To improve the accuracy and generalizability of automatic dysphonic voice detection.
Main Methods:
- A deep learning model was trained using a combination of contrastive and classification loss functions.
- Data warping techniques were applied to input voice samples to increase model robustness.
- The framework was designed to generate acoustic feature embeddings sensitive to voice quality.
Main Results:
- The proposed method achieved high classification accuracy both within and across different corpora.
- The generated embeddings demonstrated sensitivity to voice quality and robustness across varied datasets.
- The model consistently outperformed three baseline methods on clean and deteriorated voice datasets.
Conclusions:
- The developed deep learning framework offers a robust and accurate solution for automatic dysphonic voice detection.
- The method's ability to generalize across corpora makes it suitable for diverse clinical and research applications.
- The generated acoustic embeddings hold potential for further advancements in voice quality assessment.
More Related Videos
06:24Author Spotlight: Advancements in the Fabrication of Synthetic Vocal Fold Models for Phonetic and Robotic Applications
Published on: January 5, 2024
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Related Concept Videos
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Pulse amplitude and quality
A weak or absent pulse may indicate reduced cardiac output or poor left ventricular contraction, which can be signs of cardiovascular dysfunction or...
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Unsoundness of Aggregate due to Volume Change
Downsampling
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...