Related Experiment Video
Updated: Sep 22, 2025

12:43
A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis ALS
Published on: February 21, 2011
35.0K
Implementing a Statistical Parametric Speech Synthesis System for a Patient with Laryngeal Cancer
Krzysztof Szklanny1, Jakub Lachowicz1
1Multimedia Department, Polish-Japanese Academy of Information Technology, 02-008 Warsaw, Poland.
Sensors (Basel, Switzerland)
|May 20, 2022
Summary
This study developed a statistical parametric speech synthesis system for patients undergoing total laryngectomy. The generated synthetic voice achieved a high quality rating, improving quality of life for patients with laryngeal cancer.
Area of Science:
- Speech synthesis
- Medical acoustics
- Computational linguistics
Background:
- Total laryngectomy significantly impacts patient quality of life due to voice loss.
- Laryngeal cancer patients experience dysphonia, affecting communication and psychosocial well-being.
- Developing assistive voice technologies is crucial for post-laryngectomy rehabilitation.
Purpose of the Study:
- To create a statistical parametric speech synthesis system for a laryngeal cancer patient before total laryngectomy.
- To evaluate the quality of synthetic speech generated from the patient's pre-surgery recordings.
- To assess the potential of this technology to improve the quality of life for patients awaiting laryngectomy.
Main Methods:
- A statistical parametric speech synthesis model was trained using the Merlin repository and a Polish language corpus.
- Speech samples from a laryngeal cancer patient, recorded before surgery, were used for training.
- Auditory-perceptual MUSHRA listening tests were conducted with 25 experts to rate synthetic voice quality.
Main Results:
- The synthetic voice generated from the patient's recordings achieved a high MUSHRA score (69.4/100).
- This patient-derived synthetic voice outperformed a synthetic voice trained on a professional voice-over talent's recordings (63.63/100).
- The patient's pre-surgery voice exhibited dysphonia, confirmed by RBH scale and Acoustic Voice Quality Index (AVQI) analysis.
Conclusions:
- A statistical parametric speech synthesizer can be effectively created for patients awaiting total laryngectomy.
- This technology offers a promising solution to restore voice quality and enhance the mental well-being of patients.
- The developed system demonstrates the feasibility of generating high-quality synthetic speech tailored to individual patient needs.

