Related Experiment Video
Updated: Aug 30, 2026

A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
Published on: January 18, 2020
TextBridge-Track: Hierarchical Spatio-temporal Alignment for RGBE Tracking via CLIP's Textual Guidance
Abstract:
RGBE tracking aims to integrate modality-aligned RGB and event features to perform temporal identity association and spatial localization of targets. Existing methods do not effectively mitigate the modality gap between RGB frames and event data. In response, we develop a hierarchical spatio-temporal alignment framework tailored to RGBE tracking. Spatially, our framework introduces Global Indirect Alignment (GIA) and Local Direct Alignment (LDA). GIA establishes modality-level alignment by projecting RGB and event backbones into a unified CLIP-based semantic space. In parallel, LDA enforces target-region matching between template and search regions to ensure precise spatial alignment. By coupling the coarse modality-level semantic correspondence from GIA with the fine-grained target-region correspondence from LDA, we establish a spatial alignment hierarchy from global cross-modal semantics to local target discrimination under CLIP's textual guidance. TIMF further extends this hierarchy from spatial alignment to temporal-invariance alignment. By aligning predicted invariant features with features sampled from the same target sequence, it preserves identity consistency across frames and provides a temporal prior that is adaptively fused with current observations. Extensive evaluations on three benchmarks show state-of-the-art success rates: COESOT 67.0%, VisEvent 60.6%, and FELT 46.9%, with consistent improvements across both standard and challenging scenarios. Our source code is available at https: //github.com/RiverUp/TextBridge-Track.

