Related Experiment Video
Updated: Aug 3, 2025

12:39
A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
Published on: January 18, 2020
7.7K
Toward Robust Visual Object Tracking With Independent Target-Agnostic Detection and Effective Siamese Cross-Task
Summary
This study introduces a novel Siamese visual object tracking framework with target-agnostic detection and cross-task interaction. This approach enhances tracking accuracy, especially under severe appearance variations, outperforming existing methods.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Advanced Siamese networks excel in visual object tracking but struggle with severe appearance variations and lack cross-task synergy.
- Existing methods independently design classification and regression modules, hindering collaborative target localization.
Purpose of the Study:
- To develop a Siamese visual object tracking framework that addresses limitations in handling appearance variations and promotes cross-task interaction.
- To improve the robustness and accuracy of visual object tracking through enhanced target detection and unified multi-task learning.
Main Methods:
- Introduced a novel network with a target-agnostic object detection module to mitigate appearance variation issues.
- Developed a cross-task interaction module for consistent supervision of classification and regression branches.
- Implemented adaptive labels for more effective network training in multi-task learning.
Main Results:
- The proposed target-agnostic detection module effectively complements direct target inference.
- Cross-task interaction module demonstrated improved synergy between classification and regression branches.
- Experimental results on OTB100, UAV123, VOT2018, VOT2019, and LaSOT benchmarks show superior tracking performance compared to state-of-the-art methods.
Conclusions:
- The novel Siamese tracking framework with target-agnostic detection and cross-task interaction significantly enhances tracking performance.
- The proposed methods effectively address challenges posed by severe appearance variations and improve multi-task learning synergy.
- This work offers a more robust and accurate solution for visual object tracking applications.

