Related Experiment Video
Updated: May 28, 2026

A Photonic System for Generating Unconditional Polarization-Entangled Photons Based on Multiple Quantum Interference
Published on: September 5, 2019
OACodec: Audio attribute disentanglement via orthogonal disentanglement and mutual information minimization
Yukun Qian1, Wenjie Zhang1, Zehua Zhang1
1Department of Electronic and Information Engineering, Harbin Institute of Technology (Shenzhen), Shenzhen, Guangdong, China.
Abstract:
Neural audio codecs (NACs) based on end-to-end neural networks and vector quantization have recently achieved high-fidelity speech compression and reconstruction. However, most existing NACs learn entangled latent codes that mix speaker timbre, prosody, and phonetic information, which limits interpretability and controllability. Although several attempts introduce attribute-aware objectives, they often lack an explicit decomposition mechanism and a principled information-theoretic constraint to encourage independent factorization. We propose OACodec, a novel neural audio disentanglement codec specifically designed to learn disentangled representations of speech attributes. Unlike conventional neural audio codec systems that treat audio as undifferentiated data, OACodec introduces a multi-stage orthogonal disentanglement network that explicitly separates timbre, prosody and phonetic information. Each latent attribute is extracted via a residual separation mechanism, guided by orthogonality constraints and supervised learning. To further promote independent factorization, we employ mutual information estimators both within and across attribute components, minimizing their mutual information. Additionally, we adopt a smoothed Tchebycheff optimization strategy to achieve a Pareto-optimal balance between reconstruction fidelity and disentanglement objectives. Experimental results demonstrate the effectiveness of OACodec in both zero-shot voice conversion and speech reconstruction, where it outperforms VC baselines and FACodec in zero-shot VC tasks and achieves reconstruction quality comparable to EnCodec and HiFi-Codec while surpassing FACodec. We further demonstrate, by constructing a two-stage TTS model, that the disentangled codes effectively improve performance on the TTS task. Ablation studies further validate the contribution of each proposed module. OACodec lays a strong foundation for interpretable and controllable speech modeling, with promising implications for applications such as text-to-speech and speech-based language modeling. Audio samples are available at https://hamidun123.github.io/OACodecDemo/.
Related Concept Videos
Extraction: Partition and Distribution Coefficients
For extracting a solute from an aqueous phase into an organic...
¹³C NMR: ¹H–¹³C Decoupling
A broadband decoupling technique is used to simplify these complex, sometimes overlapping, signals. Broadband decoupling relies on a...
Law of Independent Assortment
Law of Independent Assortment
Vector Algebra: Method of Components
In many applications, the magnitudes and directions of...
¹H NMR: Interpreting Distorted and Overlapping Signals
As Δν decreases and the signals move closer, the doublets appear increasingly distorted. The intensities of the inner lines increase at the cost of those of the outer lines as the signals are slanted or...