Related Experiment Video
Updated: Sep 14, 2026

Utilizing vmTracking to Improve the Accuracy of Multi-Animal Pose Estimation in Rodent Social Behavior Studies
Published on: November 7, 2025
VLG-RMOT: Calibration-before-control for end-to-end referring multi-object tracking
Hongshen Zhao1, Ming Dai1, Fei Xie2
1School of Automation, Southeast University, Nanjing, 210096, China.
Abstract:
Referring multi-object tracking (RMOT) aims to detect and track all objects that satisfy a natural-language expression in video. End-to-end RMOT relies on object queries for detection, association, and temporal updating, so language must function as both an alignment cue and a control signal for evolving queries. Existing methods have advanced visual-linguistic fusion and query interaction; however, alignment alone does not ensure a scene-discriminative language signal as candidate objects, their appearance, and their relations change across frames. We therefore propose VLG-RMOT, a calibration-before-control framework that adapts language to the current visual scene before using it for query control. Specifically, cross-modal semantic calibration performs bidirectional visual-linguistic calibration with learnable channel-wise residual scaling. The resulting scene-conditioned language guides two complementary stages: contextual query synthesis establishes a shared referring prior before visual decoding, while text-aware gating modulates query states after visual cross-attention. On Refer-KITTI, Refer-KITTI-V2, and Refer-BDD, VLG-RMOT achieves the best HOTA among the compared end-to-end methods. Under a matched ablation protocol, learnable calibration improves HOTA by 1.97, 1.40, and 2.98 points over raw-language control, an unscaled residual, and fixed scaling, respectively; scaling-vector statistics and token-level visualization provide representation-level evidence of channel-selective and scene-dependent adaptation.
Related Concept Videos
Relative Motion Analysis using Rotating Axes-Problem Solving
Here, in order to determine the magnitude of velocity and acceleration for point...
Relative Motion Analysis using Rotating Axes
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it instrumental in...