Related Experiment Video
Updated: Jan 11, 2026

11:18
Closed-loop Neuro-robotic Experiments to Test Computational Properties of Neuronal Networks
Published on: March 2, 2015
10.8K
Unsupervised speech recognition through spike-timing-dependent plasticity in a convolutional spiking neural network.
Meng Dong1,2, Xuhui Huang2, Bo Xu2,3,4
1School of Automation, Harbin University of Science and Technology, Harbin, Heilongjiang, China.
Plos One
|November 30, 2018
Summary
This study introduces a biologically inspired spiking neural network (SNN) for speech recognition (SR). The novel convolutional SNN model achieves high accuracy comparable to artificial neural networks (ANNs) while offering greater biological plausibility and efficiency.
Area of Science:
- Computational Neuroscience
- Artificial Intelligence
- Signal Processing
Background:
- Artificial neural networks (ANNs) have advanced speech recognition (SR) but suffer from biological implausibility and high power consumption.
- Spiking neural networks (SNNs) offer potential solutions with efficient spike communication and brain-like synaptic plasticity for weight modification.
- Existing SNN models for SR often exhibit suboptimal performance or rely on biologically implausible training methods.
Purpose of the Study:
- To develop a biologically inspired convolutional SNN model for accurate and efficient speech recognition.
- To address the limitations of ANNs in terms of biological plausibility and power consumption for SR tasks.
- To investigate the effectiveness of unsupervised learning using spike-timing-dependent plasticity (STDP) in SNNs for SR.
Main Methods:
- A convolutional SNN model employing a time-to-first-spike coding scheme for rapid information processing.
- Unsupervised weight adjustment using spike-timing-dependent plasticity (STDP) to form receptive fields in convolutional neurons.
- Implementation of local weight sharing within the convolutional structure for enhanced feature extraction of speech signals.
Main Results:
- The SNN model achieved 97.5% accuracy on the TIDIGITS dataset when evaluated with a linear support vector machine (SVM), rivaling top ANN performance.
- Network outputs demonstrated increased linear separability, reduced dimensionality, and sparsity.
- High performance of 93.8% was obtained on the more challenging TIMIT dataset, with a linear spike-based classifier (tempotron) also achieving near-SVM accuracy.
Conclusions:
- An STDP-based convolutional SNN model, incorporating local weight sharing and temporal coding, effectively and efficiently performs speech recognition.
- The proposed SNN architecture offers a biologically plausible and power-efficient alternative to traditional ANNs for SR.
- The findings suggest that SNNs hold significant promise for advancing the field of speech recognition with brain-inspired computational principles.

