Related Experiment Video
Updated: Aug 9, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Manipulating Voice Attributes by Adversarial Learning of Structured Disentangled Representations.
Laurent Benaroya1, Nicolas Obin1, Axel Roebel1
1Analysis/Synthesis Team-STMS, IRCAM, Sorbonne University, CNRS, French Ministry of Culture, 75004 Paris, France.
This study introduces a novel neural network for voice conversion (VC) that can alter voice attributes like gender and age, not just identity. The system effectively disentangles and manipulates these features, enabling realistic voice attribute modification.
Area of Science:
- Artificial Intelligence
- Speech Processing
- Machine Learning
Background:
- Neural voice conversion (VC) has advanced significantly, enabling realistic voice identity manipulation with minimal data.
- Existing methods primarily focus on altering voice identity, leaving manipulation of other voice attributes less explored.
Purpose of the Study:
- To present an original neural architecture for voice conversion that enables manipulation of voice attributes such as gender and age.
- To disentangle speech signal information into interpretable voice attributes for flexible manipulation.
Main Methods:
- The proposed architecture is inspired by fader networks, utilizing adversarial loss to achieve mutual independence of encoded information.
- Speech signal information is disentangled into voice attributes, allowing for manipulation during inference.
- The method was evaluated on voice gender conversion using the VCTK dataset.
Main Results:
- Quantitative analysis confirmed the architecture learns gender-independent speaker representations.
- Speaker identity recognition remained accurate from the gender-independent representation.
- Subjective experiments demonstrated high efficiency and naturalness in voice gender conversion.
Conclusions:
- The novel neural architecture successfully disentangles and manipulates voice attributes, extending beyond identity conversion.
- The method achieves effective and natural-sounding voice gender conversion.
- This work opens new possibilities for nuanced voice manipulation in speech processing.
Related Concept Videos
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
Associative Learning
Classical conditioning, also known...
Larynx
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids,...
Elaborative Rehearsals
The effectiveness of...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Purposive Learning

