Related Experiment Video
Updated: Aug 13, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
LiteSwitchCodec: Neural speech coding with token space quantization and causal U-Net for personalized real-time
Xusheng Yang1, Wei Xiao2, Zixiang Wan1
1Advanced Data and Signal Processing Laboratory, School of Electronic and Computer Engineering, Peking University, Shenzhen, China.
LiteSwitchCodec offers high-quality speech compression and personalized real-time communication (PRTC) with enhanced security. This lightweight neural speech codec significantly reduces model parameters while maintaining excellent audio quality and low latency for PRTC services.
Area of Science:
- Speech Processing and Communication Systems
- Machine Learning for Audio Signal Processing
- Cybersecurity in Communication Networks
Background:
- Personalized real-time communication (PRTC) demands low latency, complexity, and robust security.
- Existing speech codecs may not adequately address the unique requirements of PRTC, including voice adaptation and copyright protection.
- The integration of voice adaptation (VA) into communication systems presents challenges in maintaining speech quality and security.
Purpose of the Study:
- To introduce LiteSwitchCodec, a novel lightweight neural speech codec designed for PRTC services.
- To enable high-quality speech compression and facilitate personalized voice adaptation within operator-managed platforms.
- To enhance security by protecting voice copyrights and preventing unauthorized impersonation in real-time communication.
Main Methods:
- Developed LiteSpeechCodec using fully causal convolutional layers for encoder/decoder with a mirrored structure for reduced decoder complexity.
- Implemented scalar quantization (SQ) as an alternative to residual vector quantization (RVQ) to simplify model architecture and improve compression.
- Introduced a lightweight causal U-Net model in the token space for personalized voice adaptation, trained using a two-stage strategy on public and custom voice conversion datasets.
Main Results:
- LiteSwitchCodec achieved a 38x reduction in model parameters compared to state-of-the-art codecs, maintaining real-time communication (RTC) quality.
- The codec demonstrated objective quality (ViSQOL: 4.32 at 7.2 kbps) and real-time voice adaptation with low latency (40 ms).
- Achieved superior performance in subjective and objective metrics for voice conversion in RTC, with 92.61% Resemblyzer similarity and a high speaker-similarity score (S-MOS: 4.67).
Conclusions:
- LiteSwitchCodec effectively balances high-quality speech compression, real-time voice adaptation, and low latency for PRTC applications.
- The codec inherently addresses security concerns, safeguarding voice copyrights and preventing impersonation, crucial for commercial RTC deployments.
- This approach offers a viable solution for next-generation communication systems prioritizing personalized experiences and robust security.
Related Concept Videos
Neuronal Communication
Signal and System
Neural Circuits
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...
Language and Cognition
Neurons as Communicators of the Brain
Cell Body
The cell body, also known...
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...