Related Experiment Videos
Adaptive multimodal visual tracking via parameter-efficient vision transformers
Yixin Xu1, Wenkang Zhang2, Tianyang Xu3
1School of Automation, Southeast University, Nanjing, 210096, China.
Abstract:
Multimodal visual object tracking (MVOT) is crucial for achieving robust performance in complex environments, including scenarios with occlusion, low illumination, or high-speed motion. However, current methods often fall short in two aspects: their reliance on static fusion strategies limits adaptability to varying scene conditions, and their use of redundant, modality-specific branches increases computational and training cost. In response, we propose SwitchTrack, a unified framework for multimodal tracking with efficient dynamic integration and on-demand modality switching. At its core lies a lightweight Dynamic Bridge module that learns content-aware fusion weights and facilitates effective interaction between modalities, adding only 0.46M parameters. To further enhance adaptability, we introduce an Task-Selective Low-Rank Adaptation strategy that enables efficient and lightweight tuning across RGB-D, RGB-T, and RGB-E configurations. This mechanism allows modality switching by updating only a small subset of parameters, without retraining the entire model. Extensive experiments on five public benchmarks demonstrate that our method achieves superior tracking accuracy and efficiency, with significantly lower training and adaptation overhead compared to state-of-the-art approaches.