Related Experiment Video
Updated: Oct 26, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.7K
Human EEG and Recurrent Neural Networks Exhibit Common Temporal Dynamics During Speech Recognition
Saeedeh Hashemnia1, Lukas Grasse1, Shweta Soni1
1Canadian Centre for Behavioural Neuroscience, Department of Neuroscience, University of Lethbridge, Lethbridge, AB, Canada.
Frontiers in Systems Neuroscience
|July 26, 2021
Summary
Deep learning models for speech recognition show similar temporal dynamics to human brain activity. Recurrent neural networks (RNNs) capture speech features like the human brain, evidenced by envelope phase tracking in EEG and RNNs.
Area of Science:
- Computational Neuroscience
- Artificial Intelligence
- Speech Processing
Background:
- Deep learning artificial neural networks excel at speech recognition, but the underlying mechanisms are not fully understood.
- Time-dependent features are crucial for speech perception in the human brain, studied via electroencephalography (EEG) and magnetoencephalography (MEG).
- Recurrent neural networks (RNNs) may emulate cortical dynamics through their temporal processing capabilities.
Purpose of the Study:
- To investigate commonalities in temporal dynamics between deep learning models and human brain activity during speech perception.
- To determine if RNNs exhibit envelope phase tracking, a phenomenon observed in human EEG.
- To compare representational similarity between RNNs and EEG signals for speech stimuli.
Main Methods:
- Presented identical sentences to human listeners and a Deep Speech RNN.
- Analyzed temporal dynamics of EEG signals and RNN units.
- Computed representational distance matrices (RDMs) for both brain and network responses.
Main Results:
- Both RNN hidden layers and human EEG signals demonstrated envelope phase tracking with comparable time lags.
- Representational distance matrices showed increasing similarity between RNN and EEG representations from early to later network layers.
- Similarity peaked at the recurrent layer of the Deep Speech RNN.
Conclusions:
- Deep Speech RNNs capture temporal speech features in a manner analogous to human brain processing.
- The findings suggest RNNs may partially emulate human cortical dynamics in speech perception.
- This research bridges artificial intelligence and neuroscience by highlighting shared temporal processing principles.

