Related Experiment Video
Updated: May 28, 2026

Functional Magnetic Resonance Imaging (fMRI) with Auditory Stimulation in Songbirds
Published on: June 3, 2013
Speech Recognition with an fMRISNN Constrained by Human Functional Brain Networks: A Study of Enhanced MFCC-Driven
Lei Guo1,2, Nancheng Ma1, Zhuoxuan Wang1
1Tianjin Key Laboratory of Bioelectromagnetic Technology and Intelligent Health, School of Health Sciences and Biomedical Engineering, Hebei University of Technology, Tianjin 300131, China.
None:
Spiking neural networks (SNNs) offer inherent advantages in processing temporal information. However, their network topologies are predominantly algorithm-generated, lacking constraints from biological brain connectivity, which limits their bio-plausibility. In our previous work, we constructed a spiking neural network (SNN) by incorporating the topological structure of functional brain networks derived from fMRI data of healthy subjects and proposed an fMRISNN model. This model was further employed as the reservoir layer of a liquid state machine (LSM) to build a speech recognition framework. In this framework, the Lyon ear model and the BSA were used to encode speech signals into spike sequences; however, this approach suffers from high computational cost and limited adaptability to temporal variations. To address these limitations, we propose an enhanced Mel-frequency cepstral coefficient (MFCC)-driven sparse spike encoding method. For the speech recognition task, we systematically compare the two preprocessing pipelines in terms of spike number, spike sparsity, encoding time, and downstream speech recognition performance. Experimental results show that the proposed method generates substantially fewer spikes, achieves markedly higher sparsity, and requires significantly less encoding time, while maintaining nearly the same recognition accuracy under the same LSM-based framework. These findings indicate that improved speech input representation can enhance the computational efficiency of SNN-based speech recognition without compromising recognition capability. In addition, the fMRISNN model significantly outperforms several baseline models with algorithmically generated topologies. Compared with mainstream models reported in the literature, although the deep convolutional neural network (CNN) still achieves higher absolute recognition accuracy, the fMRISNN exhibits clear advantages in terms of model parameter size and theoretical energy efficiency.

