Not all regions are equal: Spatially adaptive representation learning for efficient visual object tracking

Zicheng Zhang1, Shan Lin1, Hongke Xu1

  • 1The School of Electronics and Control Engineering, Chang'an University, Xi'an, 710000, Shaanxi Province, China.

Summary

This study introduces the sparse mask Transformer (SMTransformer) for visual object tracking. It enhances efficiency by adaptively learning representations and reducing redundant computations, achieving state-of-the-art accuracy and speed.