Related Experiment Video
Updated: Aug 16, 2025

06:02
Topographical Estimation of Visual Population Receptive Fields by fMRI
Published on: February 3, 2015
9.3K
A Mixed Visual Encoding Model Based on the Larger-Scale Receptive Field for Human Brain Activity
Shuxiao Ma1, Linyuan Wang1, Panpan Chen1
1Henan Key Laboratory of Imaging and Intelligent Processing, PLA Strategic Support Force Information, Engineering University, Zhengzhou 450001, China.
Brain Sciences
|December 23, 2022
Summary
This study introduces a mixed deep learning model for brain imaging analysis, enhancing visual encoding models by combining large kernel networks with traditional convolutional neural networks. This approach improves feature extraction for better understanding visual representations in the brain.
Area of Science:
- Neuroscience
- Computer Science
- Machine Learning
Background:
- Deep neural networks, particularly Convolutional Neural Networks (CNNs), are used for visual encoding models in functional magnetic resonance imaging (fMRI).
- Standard CNNs utilize small kernel sizes (e.g., 3x3), limiting their receptive field size, which is insufficient for capturing complex visual features relevant to high-level visual cortex regions.
- Biological studies show that receptive field sizes in higher visual areas are significantly larger than in lower visual areas, suggesting a need for models with larger receptive fields.
Purpose of the Study:
- To address the limitations of small receptive fields in CNNs for visual encoding models.
- To propose a novel mixed model that integrates the RepLKNet architecture with VGG for enhanced feature extraction.
- To investigate whether a larger receptive field size improves encoding performance in visual cortex regions.
Main Methods:
- Developed a mixed model by combining RepLKNet, which features large convolution kernels, with the VGG network.
- Utilized this mixed model to replace traditional CNNs for feature extraction in visual encoding tasks.
- Evaluated the model's performance across multiple regions of the visual cortex using fMRI data.
Main Results:
- The proposed mixed model demonstrated superior encoding performance in various visual cortex regions compared to traditional convolutional models.
- The integration of RepLKNet's large receptive field capability with VGG's feature extraction led to more comprehensive image feature extraction.
- Experimental results validate the hypothesis that larger receptive fields are beneficial for visual encoding models.
Conclusions:
- A larger receptive field size is crucial for developing effective visual encoding models in neuroscience research.
- The proposed mixed model offers a promising approach to enhance the role of convolutional networks in understanding visual representations.
- Future research should consider incorporating larger receptive fields to improve the performance and biological relevance of deep learning models for brain imaging analysis.

