Related Experiment Video
Updated: Jul 15, 2026

Recording and Analyzing Multimodal Large-Scale Neuronal Ensemble Dynamics on CMOS-Integrated High-Density Microelectrode Array
Published on: March 8, 2024
Multi-scale spatio-temporal learning-based neural beamformer for multichannel speech enhancement
1College of Electronic and Information Engineering, Nanjing University of Aeronautics and Astronautics, Nanjing, Jiangsu 211106, China.
Abstract:
Time-domain neural beamformers typically extract pairwise microphone features and process them using sequential networks. While this approach can capture temporal dynamics of the input features containing spatial cues, it does not explicitly model spatial and temporal information jointly, which constrains the network's ability to learn rich spatio-temporal (ST) representations for multichannel speech enhancement. This paper proposes STBFNet, a time-domain neural beamforming network that leverages multi-scale ST feature learning. In the proposed model, local ST features are first extracted and aggregated through multi-scale convolutional operations, enabling the network to jointly capture temporal dependencies and channel-wise correlations. Simultaneously, spatial information compensation is employed to enhance global channel interactions. The resulting features are fused and further processed, where self-attention is utilized to learn long-range contextual dependencies, and a modified Conformer-style convolution is applied to strengthen ST feature integration. Through these designs, the network produces refined ST representations, leading to improved performance in time-domain beamforming. Experimental results demonstrate the superiority of STBFNet over existing multichannel filtering methods. The proposed STBFNet has a smaller model size and achieves lower latency while maintaining competitive performance.