Related Experiment Video
Updated: May 11, 2025

12:49
Transcranial Direct Current Stimulation tDCS of Wernicke's and Broca's Areas in Studies of Language Learning and Word Acquisition
Published on: July 13, 2019
16.5K
CMDF-TTS: Text-to-speech method with limited target speaker corpus.
Ye Tao1, Jiawang Liu1, Chaofeng Lu1
1School of Information Science and Technology, Qingdao University of Science and Technology, Qingdao 266061, PR China.
Summary
This study introduces a fast Text-to-Speech (TTS) method using a Statistical-based Compression Auxiliary Corpus (SCAC) algorithm. It significantly reduces training costs and time for high-quality speech synthesis with minimal target speaker data.
Area of Science:
- Speech synthesis
- Machine learning
- Digital signal processing
Background:
- End-to-end Text-to-Speech (TTS) models require large auxiliary corpora, increasing training costs.
- Existing methods struggle with high-quality synthesis using limited target speaker data.
Purpose of the Study:
- To develop a fast and high-quality speech synthesis approach with minimal target speaker recordings.
- To reduce the substantial training costs associated with large auxiliary corpora in TTS.
Main Methods:
- Proposed a Statistical-based Compression Auxiliary Corpus (SCAC) algorithm to optimize auxiliary data.
- Developed a non-autoregressive model, CMDF-TTS, incorporating multi-level prosody and Denoising Diffusion Probabilistic Models (DDPMs).
- Utilized Conditional Variational Auto-Encoder Generative Adversarial Networks (CVAE-GAN) for further quality enhancement.
Main Results:
- The SCAC algorithm significantly improves model training speed without compromising speech naturalness.
- CMDF-TTS, enhanced by SCAC, demonstrates a balance between training efficiency and synthesized speech quality.
- Experimental results show superior performance compared to state-of-the-art models on Mandarin and English datasets.
Conclusions:
- The proposed SCAC algorithm effectively reduces auxiliary corpus requirements for TTS.
- CMDF-TTS offers a promising solution for efficient and high-quality speech synthesis.
- The approach successfully integrates prosody modeling and diffusion models for enhanced speech generation.
Keywords:
Compress corpusDenoising diffusion probabilistic modelsLimited target speaker corpusMulti-level prosody modelingText-to-speech
