Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Air-entraining Agents01:27

Air-entraining Agents

115
Air-entraining agents improve the durability and workability of concrete in climates with frequent freezing and thawing. These agents prevent cracks by introducing small air bubbles into the mix, creating spaces accommodating water expansion when temperatures drop. The air-entraining agents lower the surface tension of water, forming stable, small air bubbles. This method is more effective than having accidental large voids, as the intentional, smaller, and evenly distributed air voids improve...
115
Classification of Signals01:30

Classification of Signals

996
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
996
Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

510
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
510
Force Classification01:22

Force Classification

1.8K
Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
1.8K
Masking and Demasking Agents01:19

Masking and Demasking Agents

2.7K
EDTA titrations may necessitate masking and demasking agents to temporarily protect a particular metal ion in a mixture from the EDTA reaction. These agents facilitate the sequential analysis of the metal ions by forming stable complexes with some—but not all—metal ions during certain steps.
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
2.7K
Reconstruction of Signal using Interpolation01:10

Reconstruction of Signal using Interpolation

395
Signal processing techniques are essential for accurately converting continuous signals to digital formats and vice versa. When a continuous signal is sampled with a period T, the resulting sampled signal exhibits replicas of the original spectrum in the frequency domain, spaced at intervals equal to the sampling frequency. To handle this sampled signal, a zero-order hold method can be applied, which creates a piecewise constant signal by retaining each sample's value until the next...
395

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A phospholipid Camptothecin-Niraparib conjugate self-assembled into supramolecular nanotubes for combination cancer therapy.

Journal of controlled release : official journal of the Controlled Release Society·2026
Same author

Shift work and early vascular aging: a systematic review.

BMC public health·2026
Same author

Corrigendum to "Osteogenic promotion by naringin through the PI3K/AKT/mTOR pathway-mediated activation of autophagy and inhibition of apoptosis" [Bone 2026 Jun 2:211:117958/doi:10.1016/j.bone.2026.117958].

Bone·2026
Same author

Structure-programmable solid electrolytes for stack-pressure-free all-solid-state micro-batteries.

National science review·2026
Same author

Efficacy of dupilumab for pruritus-driven trichoteiromania in an atopic background: a retrospective study.

Journal of the American Academy of Dermatology·2026
Same author

Negative-Pressure-Actuated Microfluidics: A Dual-Mode Point-of-Care Sensor for Allergen-Specific IgE in Interstitial Fluid.

Analytical chemistry·2026

Related Experiment Video

Updated: Oct 9, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.7K

VSUGAN unify voice style based on spectrogram and generated adversarial networks.

Tongjie Ouyang1, Zhijun Yang2, Huilong Xie2

  • 1Modern Educational Technology and Practice Training Center, Innovation Laboratory for Undergraduate, Xiamen University, Xiamen, China. oytj@xmu.edu.cn.

Scientific Reports
|December 22, 2021
PubMed
Summary

This study introduces a voice style unification model (VSUGAN) using generative adversarial networks to improve audio quality. VSUGAN effectively reduces style differences in recorded courses across various environments without speaker-specific retraining.

More Related Videos

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

573
Asthma Detection Research Based on Voice Signal Processing and Machine Learning
04:04

Asthma Detection Research Based on Voice Signal Processing and Machine Learning

Published on: July 22, 2025

515

Related Experiment Videos

Last Updated: Oct 9, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.7K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

573
Asthma Detection Research Based on Voice Signal Processing and Machine Learning
04:04

Asthma Detection Research Based on Voice Signal Processing and Machine Learning

Published on: July 22, 2025

515

Area of Science:

  • Audio Processing
  • Machine Learning
  • Speech Synthesis

Background:

  • Audio recordings from different environments and pickups exhibit style variations.
  • These style differences negatively impact the quality of recorded courses.
  • Voice style unification is a common technique to address these issues.

Purpose of the Study:

  • To propose a novel voice style unification model named VSUGAN.
  • To enable audio style unification across diverse environments without retraining for new speakers.
  • To enhance the quality of recorded audio by reducing style discrepancies.

Main Methods:

  • Developed a voice style unification model based on generative adversarial networks (VSUGAN).
  • The VSUGAN model transfers voice style from spectrograms.
  • Audio synthesis is achieved by merging style information from an audio style template and voice information from processed audio.

Main Results:

  • VSUGAN was implemented and evaluated on the THCHS-30 and VCTK-Corpus datasets.
  • The model effectively synthesizes audio, unifying voice style.
  • Demonstrated significant improvement in recorded audio quality and reduction of style differences.

Conclusions:

  • VSUGAN effectively improves recorded audio quality.
  • The model successfully reduces style differences in audio recorded in various environments.
  • VSUGAN offers a flexible solution for audio style unification without requiring speaker-specific network retraining.