Summary

This study introduces LaVST for language-supervised multi-camera multi-object tracking (MC-MOT). It achieves competitive performance against identity-supervised methods, highlighting the potential of language annotations in MC-MOT.