Related Experiment Video
Updated: Sep 15, 2025

A Lightweight, Headphones-based System for Manipulating Auditory Feedback in Songbirds
Published on: November 26, 2012
Lightweight real-time speech enhancement: State-space models and multi-spectral scanning techniques
Xiaodong Zhu1, Junqi Yang1, Yuhong Yang1
1National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University, China; Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University, China.
Abstract:
Achieving efficient, real-time speech enhancement requires a careful balance between signal quality and computational complexity. In this paper, we propose a lightweight end-to-end framework that leverages Mamba State-Space Model (SSM) in combination with multi-spectral scanning techniques to improve speech signals in real-time communication systems. To address the distinct challenges of speech processing, we introduce three novel spectrogram scanning methods designed to capture long-term dependencies across full-band, sub-band, and cross-band spectrums. These techniques are enhanced by ERB compression, which simulates human auditory perception, and high-frequency reconstruction to reduce computational overhead. Additionally, we incorporate a Voice Activity Detection (VAD) loss function to refine the recovery of coarse-grained temporal features during training. Our model is trained using a large, synthetic dataset generated through a custom pipeline and evaluated on the ICASSP SSI Challenge dataset. The experimental results demonstrate that our approach outperforms state-of-the-art models in terms of overall quality (OVRL) and speech signal clarity (SIG), while remaining highly efficient in terms of resource consumption. Besides, on the VoiceBank+DEMAND benchmark, our framework surpasses competitive models in overall speech enhancement, demonstrating the potential of our lightweight solution for practical applications.

