Related Experiment Video
Updated: Jun 11, 2025

10:16
Synthetic, Multi-Layer, Self-Oscillating Vocal Fold Model Fabrication
Published on: December 2, 2011
14.0K
HiddenSinger: High-quality singing voice synthesis via neural audio codec and latent diffusion models
Ji-Sang Hwang1, Sang-Hoon Lee2, Seong-Whan Lee1
1Department of Artificial Intelligence, Korea University, 02841, Seoul, Republic of Korea.
Summary
HiddenSinger utilizes a neural audio codec and latent diffusion models to generate high-quality singing voices. This novel approach overcomes limitations in current singing voice synthesis (SVS) systems, even when trained on unlabeled data.
Area of Science:
- Artificial Intelligence
- Speech Synthesis
- Machine Learning
Background:
- Denoising diffusion models show promise in generative tasks but face challenges in complex, time-varying audio synthesis.
- Singing voice synthesis (SVS) requires high-dimensional samples and long-term acoustic features, posing complexity issues for existing diffusion models.
Purpose of the Study:
- To propose HiddenSinger, a high-quality SVS system addressing complexity and controllability limitations.
- To introduce an unsupervised learning framework, HiddenSinger-U, for training SVS models with unlabeled data.
Main Methods:
- A neural audio codec with an autoencoder compresses audio into a low-dimensional latent vector for high-fidelity reconstruction.
- Latent diffusion models are employed to sample latent representations from musical scores.
- The HiddenSinger-U framework enables training using unlabeled singing voice datasets.
Main Results:
- HiddenSinger demonstrates superior audio quality compared to previous SVS models.
- HiddenSinger-U successfully synthesizes high-quality singing voices trained exclusively on unlabeled data.
Conclusions:
- HiddenSinger offers a robust solution for high-quality, controllable singing voice synthesis.
- The unsupervised HiddenSinger-U framework expands the applicability of SVS models by leveraging unlabeled datasets.
Keywords:
Generative modelLatent diffusion modelNeural audio codecSinging voice synthesisUnsupervised learningMore Related Videos
Related Concept Videos
Downsampling
134
When considering a sampled sequence with zero values between sampling instants, one can replace it by taking every N-th value of the sequence. At these integer multiples of N, the original and sampled sequences coincide. This process, known as decimation, involves extracting every N-th sample from a sequence, thereby creating a more efficient sequence.
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
134
Perceiving Loudness, Pitch, and Location
198
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
198
Auditory Perception
322
The auditory system is essential for sound perception, utilizing various critical structures. When sound waves enter the outer ear, they travel through the ear canal and cause the eardrum to vibrate. These vibrations are then transmitted to the middle ear, where three tiny bones – the malleus, incus, and stapes – amplify the sound. This amplification is crucial, as it ensures that the sound vibrations are strong enough to be conveyed to the inner ear. These vibrations then reach the...
322
Larynx
1.3K
The human larynx, often referred to as the voice box, is an intricate organ located in the neck. It serves as a pathway for air to enter the lungs during respiration and is an essential component of voice production.
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids,...
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids,...
1.3K
Pulse amplitude and quality
1.7K
Pulse amplitude is a crucial indicator of cardiac health because it provides valuable insights into the strength of left ventricular contractions and the overall uniformity of blood circulation within the vasculature. The strength of the pulse is directly related to the force with which the heart contracts and the volume of blood being pumped.
A weak or absent pulse may indicate reduced cardiac output or poor left ventricular contraction, which can be signs of cardiovascular dysfunction or...
A weak or absent pulse may indicate reduced cardiac output or poor left ventricular contraction, which can be signs of cardiovascular dysfunction or...
1.7K
Auditory Pathway
5.3K
Auditory pathways constitute the complex neural circuits responsible for transmitting and interpreting auditory information from the peripheral auditory system to the brain. Sound waves are initially captured by the outer ear, funneled through the ear canal, and reach the tympanic membrane (eardrum). These vibrations are transmitted via the middle ear's ossicles to the inner ear's cochlea.
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...
5.3K

