Related Experiment Video
Updated: Oct 2, 2026

Co-analysis of Brain Structure and Function using fMRI and Diffusion-weighted Imaging
Published on: November 8, 2012
DeID-Aligner: Fine-Grained Brain-Vision Semantic Alignment for Cross-Subject Visual Decoding from Limited fMRI Data
Xu Yin1, John Q Gan2, Ming Yang3
1Department of Anesthesia, Affiliated Hospital of Xuzhou Medical University, Xuzhou 221002, Jiangsu, PR China; Key Laboratory of Child Development and Learning Science of Ministry of Education, School of Biological Science & Medical Engineering, Southeast University, Nanjing 211189, Jiangsu, PR China.
Abstract:
Recent advances in fMRI-based semantic decoding and visual reconstruction have achieved remarkable progress. However, severe inter-subject variability and the scarcity of paired fMRI-image data remain major obstacles to cross-subject generalization. To overcome these challenges, we propose DeID-Aligner, a brain functional alignment framework that enables multi-task visual decoding across subjects under limited fMRI data. The method first constructs a region-of-interest (ROI)-wise fMRI encoder that captures regional functional representations. A dynamic Patch-to-ROI fusion mechanism with Progressive Modal Attention Transition (PMAT) explicitly associates image patches with cortical regions, enabling fine-grained cross-modal semantic alignment while gradually shifting from cross-modal to fMRI-only for robust inference. To remove subject-specific information, we introduce a de-identification task implemented with a graph attention network (DIT-GAT), which models inter-ROI interaction patterns to reduce subject-specific information and suppresses identity cues. Extensive experiments on the Natural Scenes Dataset (NSD) demonstrate that DeID-Aligner significantly outperforms state-of-the-art methods in cross-subject category classification, multi-label classification, image captioning, and image reconstruction. Voxel-wise unique variance analysis further reveals that the Patch-to-ROI fusion mechanism substantially enhances semantic representations in high-level visual areas while maintaining structural fidelity in early visual cortices. These results suggest that DeID-Aligner effectively bridges inter-subject functional disparities and provides a unified, interpretable framework for multi-task visual neural decoding.
