Related Experiment Video
Updated: Sep 13, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Recurrent neural networks as neuro-computational models of human speech recognition
Christian Brodbeck1,2, Thomas Hannagan2,3, James S Magnuson2,4,5
1Department of Computing and Software, McMaster University, Hamilton, Ontario, Canada.
Recurrent Neural Networks (RNNs), specifically Long Short-Term Memory (LSTM) models, can computationally mimic human speech recognition. Their internal dynamics predict neural responses, suggesting RNNs are viable models for cortical speech processing.
Area of Science:
- Computational Neuroscience
- Cognitive Science
- Artificial Intelligence
Background:
- Human speech recognition involves processing continuous acoustic signals into discrete linguistic units.
- Recurrent Neural Networks (RNNs) process sequential data, making them potential models for temporal information processing in speech.
- Previous research suggests RNNs can simulate behavioral aspects of speech and language processing, but their neural dynamics remain unexamined.
Purpose of the Study:
- To investigate whether the internal computational dynamics of RNNs, when trained for speech recognition, resemble human neural speech processing.
- To determine if RNNs can predict human neural population responses to auditory stimuli.
- To explore how cognitive principles in RNN architecture influence their predictive power for human neural activity.
Main Methods:
- Trained Long Short-Term Memory (LSTM) RNNs on auditory spectrograms for speech recognition tasks.
- Compared the predictive accuracy of RNN internal dynamics against auditory features for human neural responses.
- Modified RNN architectures based on cognitive principles, such as phonetic competition, to assess impact on temporal dynamics.
Main Results:
- The internal dynamics of LSTM RNNs trained for speech recognition significantly predicted human neural population responses.
- RNN predictions surpassed those based solely on auditory features, indicating a deeper level of processing.
- Architectural modifications inspired by human phonetic competition enhanced the RNNs' ability to model human-like temporal dynamics.
Conclusions:
- RNNs, particularly LSTMs, offer plausible computational models for the cortical mechanisms underlying human speech recognition.
- The study demonstrates that RNN internal dynamics can mirror neural processes in the human brain during speech perception.
- Cognitive principles can be integrated into RNNs to improve their fidelity as models of human neural computation.
Related Concept Videos
Neural Circuits
Neuronal pools are collections of nerve cells with similar functions and interact through chemical and electrical signals. These pools include both interneurons (the central neural circuit nodes that...
Neurons as Communicators of the Brain
Cell Body
The cell body, also known...
Neural Regulation
Neuroplasticity
Neural Control of Respiration
Respiratory Centers in the Brainstem
Two primary areas comprise the respiratory center: the medullary respiratory center in the medulla oblongata and the pontine respiratory group in the pons. The...
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...

