Harnessing Knowledge From Pretrained VLMs for Unsupervised Person Search

Summary

This study introduces FMUPS, a new unsupervised person search method using semantic information from vision-language models (VLMs) to create reliable pseudo-labels. It overcomes challenges in generating accurate bounding boxes and identities for better pedestrian detection and re-identification.

Related Concept Videos