Related Experiment Video
Updated: Aug 13, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
LiteSwitchCodec: Neural speech coding with token space quantization and causal U-Net for personalized real-time
Xusheng Yang1, Wei Xiao2, Zixiang Wan1
1Advanced Data and Signal Processing Laboratory, School of Electronic and Computer Engineering, Peking University, Shenzhen, China.
Abstract:
With the rapid development of online conferences and live streaming, personalized real-time communication (PRTC) has emerged as a critical capability for next-generation communication systems, placing demands on latency, complexity, and security. This paper presents LiteSwitchCodec, a lightweight neural speech codec specifically designed for PRTC services on operator-managed platforms. It aims to achieve high-quality speech compression while facilitating PRTC by integrating a voice adaptation (VA) module, thereby avoiding potential risks to voice copyright and security. For speech compression, we first design LiteSpeechCodec, which employs fully causal convolutional layers as the encoder and decoder and reduces the complexity of the decoder through a mirrored structure. We introduce scalar quantization (SQ) as an alternative to residual vector quantization (RVQ), reducing model complexity while maintaining high speech quality. This approach facilitates the learning of a high-quality compression domain, which in turn simplifies the generation of the quantized tokens. A lightweight causal U-Net model is introduced in the token space to extract global information for personalized VA, supporting dynamic switching of target speakers. Specifically, we propose a two-stage training strategy. First, we train the LiteSpeechCodec on public datasets for speech compression. We then construct a voice conversion (VC) dataset to train the token-level causal U-Net VA network. Experiments demonstrate that LiteSwitchCodec maintains RTC quality while reducing model parameters by 38 × compared to the state-of-the-art codec, achieving an objective quality of ViSQOL: 4.32 at 7.2 kbps. Moreover, LiteSwitchCodec achieves real-time VA with a low latency of 40 ms. Compared with VC models in RTC transmission, our method shows superior performance in both subjective and objective metrics, achieving an objective Resemblyzer similarity of 92.61% and a subjective speaker-similarity score (S-MOS: 4.67 vs. 3.59), highlighting its effectiveness in PRTC. Critically, LiteSwitchCodec inherently safeguards voice copyrights and prevents unauthorized impersonation, satisfying the security demands of commercial RTC deployments.
Related Concept Videos
Neuronal Communication
Signal and System
Neural Circuits
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...
Language and Cognition
Neurons as Communicators of the Brain
Cell Body
The cell body, also known...
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...