VGSG: Vision-Guided Semantic-Group Network for Text-Based Person Search

Summary

This study introduces a Vision-Guided Semantic-Group Network (VGSG) for efficient text-based person search. The VGSG network effectively aligns fine-grained visual and textual features without external tools, improving retrieval accuracy.