Related Experiment Video
Updated: Sep 16, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Transinger: Cross-Lingual Singing Voice Synthesis via IPA-Based Phonetic Alignment
Chen Shen1, Lu Zhao1, Cejin Fu1
1College of Computer and Information Engineering (College of Artificial Intelligence), Nanjing Tech University, Nanjing 211816, China.
Transinger enhances singing voice synthesis (SVS) by using International Phonetic Alphabet (IPA) for cross-lingual generalization. This novel approach improves pronunciation modeling for unseen languages, advancing SVS capabilities.
Area of Science:
- Artificial Intelligence
- Speech Technology
- Computational Linguistics
Background:
- Singing Voice Synthesis (SVS) faces challenges with global linguistic diversity.
- Fragmented, language-specific phoneme encodings hinder unified phonetic modeling in SVS.
- Existing SVS research lacks exploration of cross-lingual generalization.
Purpose of the Study:
- To develop a cross-lingual singing voice synthesis framework that overcomes linguistic barriers.
- To improve phonetic representation and generalization capabilities in SVS models.
- To enable SVS for languages not explicitly included in the training data.
Main Methods:
- Created a four-language dataset using International Phonetic Alphabet (IPA) for consistent phonetic representation.
- Proposed a novel method of decomposing IPA phonemes into letters and diacritics for deeper pronunciation rule learning.
- Introduced Transinger, a cross-lingual synthesis framework integrating Conformer and RVQ techniques.
- Developed a dynamic IPA adaptation strategy for applying learned representations to unseen languages.
Main Results:
- Transinger demonstrated significant improvements in cross-lingual generalization compared to state-of-the-art methods.
- Objective and subjective experiments confirmed the effectiveness of the proposed approach.
- The method achieved outstanding cross-lingual synthesis performance, even for languages not seen during training.
Conclusions:
- Multilingual aligned representations enhance SVS model learning efficacy and robustness.
- Decomposing IPA phonemes into letters and diacritics improves pronunciation learning and generalization.
- Transinger represents a breakthrough in cross-lingual singing voice synthesis.
Related Concept Videos
Larynx
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids,...
Improving Translational Accuracy
Pharynx
Nasopharynx
The nasopharynx, bordered by the conchae of the nasal cavity, serves exclusively as an air conduit. In its superior region, the pharyngeal tonsils or adenoids are located. These tonsils are clusters of lymphoid reticular tissue akin to a lymph node. The precise...
Air-entraining Agents
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....

