Related Experiment Video
Updated: Sep 15, 2025

10:13
A Lightweight, Headphones-based System for Manipulating Auditory Feedback in Songbirds
Published on: November 26, 2012
14.5K
Lightweight real-time speech enhancement: State-space models and multi-spectral scanning techniques
Xiaodong Zhu1, Junqi Yang1, Yuhong Yang1
1National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University, China; Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University, China.
Summary
This study introduces a lightweight Mamba State-Space Model (SSM) framework for real-time speech enhancement. The novel multi-spectral scanning techniques significantly improve speech quality and clarity while maintaining high efficiency.
Area of Science:
- Signal Processing
- Machine Learning
- Auditory Perception
Background:
- Real-time speech enhancement demands balancing signal quality and computational cost.
- Existing methods often struggle with long-term dependencies and computational overhead.
Purpose of the Study:
- To develop a lightweight, end-to-end speech enhancement framework for real-time communication.
- To improve speech signal quality and clarity efficiently.
Main Methods:
- Utilized Mamba State-Space Model (SSM) with novel multi-spectral scanning techniques (full-band, sub-band, cross-band).
- Incorporated ERB compression for auditory perception simulation and high-frequency reconstruction.
- Employed Voice Activity Detection (VAD) loss for temporal feature refinement.
- Trained on a custom synthetic dataset and evaluated on ICASSP SSI and VoiceBank+DEMAND benchmarks.
Main Results:
- The proposed framework achieved superior performance in overall quality (OVRL) and speech clarity (SIG) compared to state-of-the-art models.
- Demonstrated significant improvements on the VoiceBank+DEMAND benchmark for overall speech enhancement.
- Maintained high computational efficiency, making it suitable for practical applications.
Conclusions:
- The lightweight Mamba-SSM framework offers an effective solution for real-time speech enhancement.
- Novel multi-spectral scanning and auditory-inspired techniques contribute to enhanced performance.
- The model presents a promising, efficient approach for practical speech processing applications.

