Related Experiment Video
Updated: Aug 28, 2026

Three-Dimensional Kinematic Characterization of an Object Pick-Up Task Using Motion Capture
Published on: May 15, 2026
CLIP-SGI: A Semantic-Guided and Instance-Consistent Framework for Generalizable Person Re-Identification
Abstract:
Generalizable person re-identification (ReID) requires a model trained on labeled source domains to remain discriminative in unseen environments. Although vision-language models provide rich cross-modal priors, the appearance semantics used by CLIP-based ReID are often encoded implicitly in learned prompts and are not explicitly organized into reusable part-level cues. Moreover, conventional identity supervision mainly emphasizes the separation of source identities and makes limited use of the local relationships among visually similar instances. To address these issues, we propose CLIP-SGI, a semantic-guided and instance-consistent framework for generalizable person ReID. First, multiple off-the-shelf vision-language models generate pedestrian descriptions. For each VLM, upper- and lower-body attributes are first voted across images of the same identity and then across the VLM-specific identity labels. Second, we construct an Attribute Prototype Bank (APB) that uses the resulting attributes as region-aware semantic anchors to guide appearance-sensitive feature learning. Third, we introduce a similarity-aware and frequency-normalized soft-label constraint that preserves ground-truth identity supervision while exploiting reliable neighborhood relationships as auxiliary signals. The three-stage training scheme combines semantic guidance, domain-aware representation learning, and instance consistency to improve robustness under domain shifts. Extensive experiments on multiple benchmark datasets demonstrate the effectiveness of the proposed method and its consistent improvements in mean average precision (mAP) and Rank-1 (R1) accuracy.
Related Concept Videos
Social Identity
Understanding Self-Concept