Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Distributed quantum sensing with measurement-after-interaction strategies.

NPJ quantum information·2026
Same author

Erratum to "Remnant cholesterol inflammatory index and MASLD in U.S. adults: mediation role of triglyceride-glucose index". [Endocrinology 43 (2026) 100427.

Journal of clinical & translational endocrinology·2026
Same author

Finite Element Analysis of Upper Airway in Ansa Cervicalis Stimulation for Obstructive Sleep Apnea.

The Laryngoscope·2026
Same author

Relaunching arcade games: game nostalgia and collective memory practice for China's first-generation players.

Frontiers in psychology·2026
Same author

Application of GBD 2021 and multiple omics to reveal the association and potential mechanism between glycated hemoglobin and intervertebral disc degeneration.

Naunyn-Schmiedeberg's archives of pharmacology·2026
Same author

Oral microbiome and Frailty: Insights from NHANES 2009-2012 and Mendelian Randomization Analysis.

The journals of gerontology. Series A, Biological sciences and medical sciences·2026

Related Experiment Video

Updated: Aug 28, 2025

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
04:04

Asthma Detection Research Based on Voice Signal Processing and Machine Learning

Published on: July 22, 2025

350

Deep learning in automatic detection of dysphonia: Comparing acoustic features and developing a generalizable

Zhen Chen1,2, Peixi Zhu3, Wei Qiu4

  • 1Department of Rehabilitation Sciences, East China Normal University, Shanghai, China.

International Journal of Language & Communication Disorders
|September 19, 2022
PubMed
Summary

Deep learning (DL) effectively identifies dysphonia using mel-spectrograms, outperforming traditional methods. This AI framework offers a reliable approach for vocal health screening and automatic voice assessment.

Keywords:
convolutional neural networkdeep learningdysphoniamel-frequency cepstral coefficientsmel-spectrogram

More Related Videos

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

519
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.6K

Related Experiment Videos

Last Updated: Aug 28, 2025

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
04:04

Asthma Detection Research Based on Voice Signal Processing and Machine Learning

Published on: July 22, 2025

350
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

519
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.6K

Area of Science:

  • Artificial intelligence
  • Machine learning
  • Speech science

Background:

  • Auditory-perceptual voice assessment is subjective and limited by rater reliability.
  • Deep learning (DL) offers potential for consistent and accessible voice analysis.
  • The performance of DL models on various acoustic features for dysphonia detection is not well-established.

Purpose of the Study:

  • To develop a generalizable DL framework for identifying dysphonia.
  • To evaluate the effectiveness of different acoustic features within a DL model.
  • To compare DL model performance against traditional machine learning approaches.

Main Methods:

  • Collected sustained phonation recordings from 238 dysphonic and 223 healthy Chinese Mandarin speakers.
  • Extracted Mel frequency cepstral coefficients (MFCCs) and mel-spectrograms from normalized audio segments.
  • Utilized a convolutional neural network (CNN) for binary classification and cross-validation.

Main Results:

  • The mel-spectrogram feature achieved the highest performance, with 97.2% AUC and 92% accuracy.
  • The DL framework significantly outperformed two baseline machine learning models on both Chinese and German test sets.
  • Optimal classification using all segments of both vowels yielded 95% accuracy on the Chinese set and 92% on the German set.

Conclusions:

  • DL provides a feasible method for the automatic detection of dysphonia.
  • Mel-spectrograms are a preferred acoustic feature for DL-based voice analysis.
  • The developed DL framework can be applied to vocal health screening and remote voice monitoring.