Related Experiment Video
Updated: Jun 5, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Enhancing Dysarthric Voice Conversion with Fuzzy Expectation Maximization in Diffusion Models for Phoneme Prediction
Wen-Shin Hsu1,2, Guang-Tao Lin1, Wei-Hsun Wang3,4,5,6
1Department of Medical Information, Chung Shan Medical University, Taichung 402201, Taiwan.
This study introduces a new voice conversion method using Fuzzy Expectation Maximization (FEM) and diffusion models to improve phoneme prediction for dysarthric speech. The approach enhances speech intelligibility and naturalness for individuals with motor speech disorders.
Area of Science:
- Speech Processing
- Artificial Intelligence
- Assistive Technologies
Background:
- Dysarthria, a motor speech disorder, impairs speech intelligibility due to neurological damage.
- Existing voice conversion (VC) systems struggle with the variability of dysarthric speech, particularly in phoneme prediction.
- Accurate phoneme prediction is crucial for improving dysarthric VC quality and communication.
Purpose of the Study:
- To develop a novel approach for enhanced phoneme prediction in dysarthric speech.
- To improve the quality and intelligibility of voice conversion for individuals with dysarthria.
- To integrate Fuzzy Expectation Maximization (FEM) with diffusion models for robust speech processing.
Main Methods:
- Combined Fuzzy Expectation Maximization (FEM) clustering with Diffusion Probabilistic Models (DPM).
- Utilized diffusion models for noise simulation to enhance speech signal robustness.
- Employed FEM for iterative phoneme boundary optimization to reduce uncertainty.
- Trained the system on the Saarland University Voice Disorder dataset, processing speech in the Mel-spectrogram domain.
Main Results:
- Significantly improved phoneme prediction accuracy and overall voice conversion quality.
- Achieved higher Mean Opinion Scores (MOS) for naturalness, intelligibility, and speaker similarity compared to StarGAN-VC and CycleGAN-VC.
- Demonstrated lower Word Error Rate (WER) for both mild and severe dysarthria, indicating enhanced speech intelligibility.
Conclusions:
- The integration of FEM and diffusion models substantially improves handling dysarthric speech irregularities.
- The method shows robustness, maintaining speech naturalness and intelligibility without a speaker-encoder.
- This approach offers a promising foundation for developing reliable assistive communication technologies and personalized speech therapy for dysarthria.
Related Concept Videos
Improving Translational Accuracy
Determination of Expected Frequency
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Expected Frequencies in Goodness-of-Fit Tests
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Propagation of Uncertainty from Random Error

