Related Experiment Video
Updated: Jul 9, 2025

07:36
Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
15.7K
VGSG: Vision-Guided Semantic-Group Network for Text-Based Person Search
Summary
This study introduces a Vision-Guided Semantic-Group Network (VGSG) for efficient text-based person search. The VGSG network effectively aligns fine-grained visual and textual features without external tools, improving retrieval accuracy.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Text-based Person Search (TBPS) requires aligning fine-grained visual and textual features across modalities.
- Existing TBPS methods often rely on inefficient external tools or complex cross-modal interactions for feature alignment.
Purpose of the Study:
- To propose an efficient Vision-Guided Semantic-Group Network (VGSG) for text-based person search.
- To extract well-aligned fine-grained visual and textual features without external tools or heavy cross-modal interaction.
Main Methods:
- Developed a Semantic-Group Textual Learning (SGTL) module to group textual features based on semantic cues.
- Implemented a Vision-guided Knowledge Transfer (VGKT) module using vision-guided attention for feature alignment.
- Introduced relational knowledge transfer (vision-language similarity and class probability transfer) for adaptive information propagation.
Main Results:
- The VGSG network effectively extracts aligned fine-grained visual and textual features.
- SGTL implicitly groups similar semantic patterns in textual features.
- VGKT aligns semantic-group textual features with visual features using relational knowledge transfer.
Conclusions:
- The proposed VGSG method achieves superior performance on text-based person search benchmarks.
- The approach offers an efficient and effective solution for cross-modal feature alignment in TBPS.
More Related Videos
Related Concept Videos
Prosopagnosia
175
Prosopagnosia, also known as face blindness, is the inability to recognize faces. In severe cases, individuals with prosopagnosia may not recognize close family members, including parents and spouses, by their faces. For instance, someone with prosopagnosia might walk past their child in a crowd, only realizing their mistake upon noticing their child's distinctive backpack or favorite jacket. Prosopagnosia specifically impairs facial recognition, while the recognition of other objects or...
175
Stereotype Content Model
14.7K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.7K

