VGSG:以视觉为导向的语义组网络,用于基于文本的个人搜索
概括
本研究介绍了一种视觉引导的语义组网络 (VGSG),用于高效的基于文本的人员搜索. 在没有外部工具的情况下,VGSG网络有效地调整了细粒度的视觉和文本特征,提高了检索准确性.
科学领域:
- 计算机视觉 计算机视觉
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 基于文本的人员搜索 (TBPS) 需要在各种模式中对准细粒度的视觉和文本特征.
- 现有的TBPS方法通常依赖于低效的外部工具或复杂的跨模式交互来实现特征对齐.
研究的目的:
- 提出一个高效的视觉指导语义组网络 (VGSG) 基于文本的人搜索.
- 在没有外部工具或沉重的交叉模式交互的情况下,提取精细细粒度的视觉和文本特征.
主要方法:
- 开发了一个语义组文本学习 (SGTL) 模块,以根据语义线索对文本特征进行分组.
- 实施了视觉引导知识传递 (VGKT) 模块,使用视觉引导注意力进行特征对齐.
- 引入了关系知识转移 (视觉语言相似性和类概率转移) 以适应信息传播.
主要成果:
- VGSG网络有效地提取了对齐的细粒度视觉和文本特征.
- 在文本特征中,SGTL隐含地将类似的语义模式组合在一起.
- VGKT将语义组的文本特征与使用关系知识转移的视觉特征对齐.
结论:
- 拟议的VGSG方法在基于文本的个人搜索基准上取得了卓越的性能.
- 该方法为TBPS的跨模态特征对齐提供了高效和有效的解决方案.
相关概念视频
Prosopagnosia
175
Prosopagnosia, also known as face blindness, is the inability to recognize faces. In severe cases, individuals with prosopagnosia may not recognize close family members, including parents and spouses, by their faces. For instance, someone with prosopagnosia might walk past their child in a crowd, only realizing their mistake upon noticing their child's distinctive backpack or favorite jacket. Prosopagnosia specifically impairs facial recognition, while the recognition of other objects or...
175
Stereotype Content Model
14.7K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.7K


