Related Experiment Video
Updated: Aug 6, 2025

06:04
Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
Published on: March 24, 2023
428
Quantifying and Improving the Performance of Speech Recognition Systems on Dysphonic Speech.
Julio C Hidalgo Lopez1, Shelly Sandeep1, MaKayla Wright2
1Emory University School of Medicine, Atlanta, Georgia, USA.
Summary
Current speech recognition systems struggle with dysphonic speech. A custom model significantly improved accuracy for conditions like spasmodic dysphonia, achieving over 96% performance on dysphonic voices.
Area of Science:
- Speech-language pathology
- Artificial intelligence
- Machine learning
Background:
- Current speech recognition systems exhibit performance disparities when processing dysphonic speech compared to typical speech.
- The underlying cause is hypothesized to be the underrepresentation of dysphonic voice data in the training datasets of these systems.
Purpose of the Study:
- To quantify the performance of existing speech recognition systems on dysphonic speech input.
- To investigate methods for improving speech recognition accuracy for individuals with voice disorders.
Main Methods:
- A retrospective database of dysphonic speech recordings was utilized.
- Three commercial speech recognition platforms were evaluated against both normal and dysphonic voice inputs.
- A custom speech recognition model was developed using transfer learning on data from patients with spasmodic dysphonia and vocal cord paralysis.
Main Results:
- Commercial platforms demonstrated significantly lower accuracy on dysphonic speech, with accuracy rates ranging from 84.55% to 93.56%.
- Performance deficits were particularly noted for spasmodic dysphonia and vocal fold paralysis.
- The custom-trained model achieved a high accuracy of 96.43% for dysphonic voices and 97.62% for normal voices.
Conclusions:
- Existing speech recognition technology underperforms with dysphonic speech due to limited training data.
- Transfer learning and custom model development can substantially enhance speech recognition performance for dysphonic voices, addressing specific pathologies.
More Related Videos
Related Concept Videos
Frequency-Domain Interpretation of PD Control
155
Proportional-Derivative (PD) controllers are widely used in fan control systems to improve stability and performance. A fan control system can be effectively represented using a Bode plot to illustrate the impact of a PD controller through its transfer function. The Bode plot visually conveys how PD control modifies the fan's response across various frequencies, providing a frequency domain interpretation of the controller's behavior.
The proportional control gain, combined with the...
The proportional control gain, combined with the...
155
Equipments Used To Measure Blood Pressure
1.1K
Direct Method
This invasive approach involves cannulating a peripheral artery. During each cardiac contraction, pressure generates mechanical motion within the catheter, transmitted through rigid, fluid-filled tubing to a transducer. This transducer converts mechanical motion into electrical signals displayed as waveforms on a monitor. An automatic flushing system prevents blood backflow. Due to the potential risk of unexpected arterial blood loss, this method is primarily used in intensive...
This invasive approach involves cannulating a peripheral artery. During each cardiac contraction, pressure generates mechanical motion within the catheter, transmitted through rigid, fluid-filled tubing to a transducer. This transducer converts mechanical motion into electrical signals displayed as waveforms on a monitor. An automatic flushing system prevents blood backflow. Due to the potential risk of unexpected arterial blood loss, this method is primarily used in intensive...
1.1K
Improving Translational Accuracy
11.7K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.7K

