Related Experiment Video
Updated: Sep 19, 2025

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
Learning contrastive semantic decomposition for visual grounding
Jie Wu1, Chunlei Wu1, Yiwei Wei2
1Qingdao Institute of Software, College of Computer Science and Technology, China University of Petroleum (East China), Qingdao, China.
This study introduces a new Contrastive Semantic Decomposition network for Visual Grounding (CSDVG). CSDVG improves accuracy by better separating and combining visual and language features for object identification.
Area of Science:
- Computer Vision
- Natural Language Processing
- Artificial Intelligence
Background:
- Visual grounding links natural language descriptions to image regions.
- Current methods use separate encoders, potentially missing shared attributes and causing redundant fusion.
Purpose of the Study:
- To propose a novel Contrastive Semantic Decomposition network for Visual Grounding (CSDVG).
- To effectively decompose shared-specific semantic features and model cross-modality features for improved visual grounding.
Main Methods:
- Developed CSDVG with an associated semantic branch for shared features and an independent semantic branch for specific features.
- Introduced a relevance-driven loss function to balance shared and specific feature learning.
Main Results:
- CSDVG demonstrated superior performance compared to existing approaches.
- Experiments showed effectiveness across all tested datasets.
Conclusions:
- The proposed CSDVG effectively decomposes and models cross-modality features.
- CSDVG addresses limitations of independent encoders and redundant fusion in visual grounding tasks.
More Related Videos
06:33Decomposing the Variance in Reading Comprehension to Reveal the Unique and Common Effects of Language and Decoding
Published on: October 11, 2018
09:27Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018
Related Concept Videos
Gestalt Principles of Perception
Visual System
Once through the pupil, the light passes through the lens, a...
Depth Perception and Spatial Vision
Difference from Background: Limit of Detection
The LOD indicates the presence or absence...
Deconvolution
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Structural Classification of Joints
A fibrous joint is where the adjacent bones are united by fibrous connective...