Related Experiment Video
Updated: Sep 2, 2026

End-To-End Deep Neural Network for Salient Object Detection in Complex Environments
Published on: December 15, 2023
Enhancing Neural Encoding of Natural Scenes through Hierarchical Integration of Saliency and Semantic Context
Sizhuo Wang1, Fan Qin1, Quan Pan1
1The Clinical Hospital of Chengdu Brain Science Institute, China-Cuba Belt and Road Joint Laboratory on Neurotechnology and Brain-Apparatus Communication, Brain-Computer Interface and Brain-Inspired Intelligence Key Laboratory of Sichuan Province, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu 610054, China.
Abstract:
Understanding how the human brain encodes complex natural scenes remains a central problem in computational neuroscience and artificial intelligence. Existing visual encoding models often rely on a single dominant feature representation and may insufficiently characterize how saliency-guided spatial information and high-level semantic context jointly contribute to cortical response prediction. To address this issue, this study proposes a saliency-guided multimodal visual encoding model, termed SMG-MVEM, to predict voxel-wise cortical responses to natural scene stimuli. The model integrates image features, saliency cues, and text-derived semantic representations through a hierarchical fusion architecture, followed by a Transformer-based brain mapper. Experiments on the Natural Scenes Dataset (NSD) show that SMG-MVEM improves prediction performance over representative neural encoding baselines and internal control variants, with the average PCC increasing from [Formula: see text] for the best-performing baseline to [Formula: see text]. Regional analyses further show that saliency contributed more strongly to early visual areas, whereas semantic features provided greater benefits in higher-order regions. Representational analyses also suggest that the model-predicted responses preserved aspects of hierarchical and category-related organization across the visual cortex. These findings indicate that structured integration of saliency and semantic context can improve cortical response prediction and provide interpretable representational patterns for natural vision.
Related Concept Videos
Vision
Visual System
Once through the pupil, the light passes through the lens, a...
Parallel Processing
Encoding
Automatic processing involves the encoding of details like time, space, frequency, and the meaning of words, usually done without conscious...
Gestalt Principles of Perception
Depth Perception and Spatial Vision
