Related Experiment Video
Updated: May 22, 2025

Author Spotlight: Advancements in the Fabrication of Synthetic Vocal Fold Models for Phonetic and Robotic Applications
Published on: January 5, 2024
Singing to speech conversion with generative flow
Jiawen Huang1, Emmanouil Benetos1
1Centre for Digital Music, Queen Mary University of London, London, UK.
Abstract:
This paper introduces singing to speech conversion (S2S), a cross-domain voice conversion task, and presents the first deep learning-based S2S system. S2S aims to transform singing into speech while retaining the phonetic information, reducing variations in pitch, rhythm, and timbre. Inspired by the Glow-TTS architecture, the proposed model is built using generative flow, with an adjusted alignment module between the latent features. We adapt the original monotonic alignment search (MAS) to the S2S scenario and utilize a duration predictor to deal with the duration differences between the two modalities. Subjective evaluations show that the proposed model outperforms signal processing baselines in naturalness and outperforms a transcribe-and-synthesize baseline in phonetic similarity to the original singing. We further demonstrate that singing-to-speech could be an effective augmentation method for low-resource lyrics transcription.
Related Concept Videos
Gradually Varying Flow
Introduction to Types of Flows
Two-dimensional flow involves changes in both length and height, as seen in...
Flow Sheet
Here's a closer look at the examples of flowsheets commonly used by nurses:
Graphic Sheet Documentation:
Rapidly Varying Flow
General External Flow Characteristics
Steady Flow of a Fluid Stream
During this process, the momentum of the fluid within the control volume remains constant over the time interval dt. By applying the...

