Related Experiment Video
Updated: Jan 12, 2026

Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
Incorporating multi-modal prompt learning into foundation models enhances predictability of visual fMRI responses to
Panpan Chen1, Chi Zhang1, Bao Li1
1Henan Key Laboratory of Imaging and Intelligent Processing, Information Engineering University, Zhengzhou 450000, People's Republic of China.
Abstract:
Objective. Modeling neural encoding of visual stimuli often uses deep neural networks (DNNs) to predict human brain response to external stimuli. However, each DNN depends on networks tailored for computer vision tasks, resulting in suboptimal brain correspondence. On the other hand, when end-to-end optimizing the encoding process for specific brain regions, challenges like training difficulties arise. Additionally, these models mostly focus on visual information processing, while the human brain integrates multi-modal information such as language to achieve a comprehensive understanding.Approach. To address these limitations, this paper proposes a multi-modal prompt learning (PL) model for neural encoding of dynamic natural stimuli. Specifically, we leverage the powerful representation ability of pre-trained foundation models and fine-tune them using our multi-modal prompts. These prompts, which include textual and visual prompts tailored to each specific regions of interest, can adapt foundation models to neural encoding tasks with fewer trainable parameters. We use the CLIP For video Clip retrieval (CLIP4clip) and Video Masked Autoencoder V2 (videoMAEv2) for feature extraction with backbone freezing, refine the representations via PL, and map the fused multi-modal features to predict voxel-wise brain responses.Main results. Extensive experiments on two functional magnetic resonance imaging video datasets demonstrate that our method outperforms existing fine-tuning methods and public models.Significance. This work highlights the potential of prompt-based fine-tuning strategies in bridging the gap between foundation models and neural encoding tasks.

