Related Experiment Video
Updated: Apr 20, 2026

fMRI Mapping of Brain Activity Associated with the Vocal Production of Consonant and Dissonant Intervals
Published on: May 23, 2017
A large-scale fMRI dataset for vision-language semantic association.
Shurui Li1,2,3, Zheyu Jin1,2, Shi Gu4,5
1School of Biomedical Engineering, ShanghaiTech University, Shanghai, China.
This study introduces the Caption Scene Dataset (CSD), a novel fMRI dataset linking visual and language information. Deep learning models effectively predicted neural responses, advancing vision neuroscience and artificial intelligence research.
Area of Science:
- Neuroscience
- Artificial Intelligence
- Cognitive Science
Background:
- Deep learning and large-scale brain activity datasets are crucial for understanding neural coding and visual-language associations.
- Large-scale functional magnetic resonance imaging (fMRI) datasets with naturalistic stimuli enhance ecological validity and research reproducibility.
- Existing datasets often focus on isolated modalities, limiting the study of cross-modal semantic associations.
Purpose of the Study:
- To introduce the Caption Scene Dataset (CSD), a large-scale fMRI dataset designed for investigating vision-language semantic association.
- To provide a resource for studying the neural basis of how the brain integrates visual and linguistic information.
- To facilitate interdisciplinary research bridging vision neuroscience and artificial intelligence.
Main Methods:
- Acquired fMRI neural responses from eight healthy participants viewing 4,400 pairs of Chinese captions and naturalistic scenes.
- Participants performed a task to determine semantic consistency between captions and images.
- Developed and applied deep neural encoding models to predict neural responses.
Main Results:
- Demonstrated the effectiveness of deep neural encoding models in predicting neural responses to both visual and language stimuli.
- Showcased successful prediction across various cortical regions, indicating robust neural encoding.
- Validated the utility of the CSD dataset for exploring vision-language semantic processing.
Conclusions:
- The Caption Scene Dataset (CSD) offers a valuable platform for investigating the neural underpinnings of semantic association between vision and language.
- The findings highlight the potential of deep learning models in decoding brain activity related to cross-modal semantic understanding.
- This dataset is expected to drive advancements in both neuroscience and artificial intelligence by enabling new research avenues.
More Related Videos
08:36Dynamic Inter-subject Functional Connectivity Reveals Moment-to-Moment Brain Network Configurations Driven by Continuous or Communication Paradigms
Published on: March 21, 2019
07:11Author Spotlight: Insights into Visual Cortex Research Through Wide-View fMRI Mapping
Published on: December 8, 2023