Related Experiment Video
Updated: Jan 29, 2026

Functional Imaging of Auditory Cortex in Adult Cats using High-field fMRI
Published on: February 19, 2014
A hierarchical sparse coding model predicts acoustic feature encoding in both auditory midbrain and cortex
Qingtian Zhang1, Xiaolin Hu1,2, Bo Hong3
1Department of Computer Science and Technology, Tsinghua University, Beijing, China.
Abstract:
The auditory pathway consists of multiple stages, from the cochlear nucleus to the auditory cortex. Neurons acting at different stages have different functions and exhibit different response properties. It is unclear whether these stages share a common encoding mechanism. We trained an unsupervised deep learning model consisting of alternating sparse coding and max pooling layers on cochleogram-filtered human speech. Evaluation of the response properties revealed that computing units in lower layers exhibited spectro-temporal receptive fields (STRFs) similar to those of inferior colliculus neurons measured in physiological experiments, including properties such as sound onset and termination, checkerboard pattern, and spectral motion. Units in upper layers tended to be tuned to phonetic features such as plosivity and nasality, resembling the results of field recording in human auditory cortex. Variation of the sparseness level of the units in each higher layer revealed a positive correlation between the sparseness level and the strength of phonetic feature encoding. The activities of the units in the top layer, but not other layers, correlated with the dynamics of the first two formants (F1, F2) of all phonemes, indicating the encoding of phoneme dynamics in these units. These results suggest that the principles of sparse coding and max pooling may be universal in the human auditory pathway.
More Related Videos
10:09Mapping the After-effects of Theta Burst Stimulation on the Human Auditory Cortex with Functional Imaging
Published on: September 12, 2012
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Related Concept Videos
lncRNA - Long Non-coding RNAs
lncRNA - Long Non-coding RNAs
Encoding
Automatic processing involves the encoding of details like time, space, frequency, and the meaning of words, usually done without conscious...
Predicting Molecular Geometry
Nursing Code of Ethics
The Auditory Ossicles
The aptly named stapes look very much like a stirrup. The three ossicles are unique to mammals, and each plays a role in...