Related Experiment Video
Updated: Aug 5, 2026

Utilizing vmTracking to Improve the Accuracy of Multi-Animal Pose Estimation in Rodent Social Behavior Studies
Published on: November 7, 2025
RegTrack: Simplicity Beneath Complexity in Robust Multi-Modal 3D Multi-Object Tracking
Abstract:
Existing 3D multi-object tracking (MOT) methods often trade efficiency and generalizability for robustness, as they typically rely on complex association metrics derived from multi-modal architectures or class-specific motion priors. Challenging the common belief that greater complexity necessarily leads to stronger robustness, we propose a robust, efficient, and generalizable method for multi-modal 3D MOT, dubbed RegTrack. Inspired by Yang-Mills gauge theory, RegTrack formulates multi-modal 3D MOT as motion-compensated representation learning. Under this analogy, point-cloud object representations are viewed as matter fields, while inter-frame object motions are regarded as local variations. Geometric cues are modeled as gauge fields to adaptively compensate for such variations, and a pretrained image representation space serves as a globally invariant physical law to guide the compensation process. In this way, the resulting motion-compensated point-cloud representations, viewed as observables, are encouraged to remain consistent for the same object across frames while preserving discriminability among different objects. Their pairwise similarities thus provide a simple yet robust association metric. Specifically, RegTrack is built upon a unified tri-cue encoder (UTEnc), which consists of a local-global point cloud encoder (LG-PEnc), a mixture-of-experts-based geometry encoder (MoE-GEnc), and a frozen image encoder derived from a pretrained vision-language model. LG-PEnc efficiently encodes the spatial-structural information of object point clouds to generate foundational representations. MoE-GEnc interacts with LG-PEnc to model inter-frame geometric relationships and adaptively compensate for motion-induced representation variations without relying on class-specific priors. The frozen image encoder is used only during training to provide a stable representation space for supervising the compensation process, and is discarded during inference. As a result, RegTrack achieves robust, efficient, and generalizable inference using only point-cloud inputs, with merely 2.67 M parameters. Extensive experiments on KITTI and nuScenes demonstrate that RegTrack outperforms its thirty-five competitors.
Related Concept Videos
Relative Motion Analysis using Rotating Axes
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it instrumental in...
Relative Motion Analysis using Rotating Axes-Problem Solving
Here, in order to determine the magnitude of velocity and acceleration for point...
Orthogonal Trajectories
Multi-input and Multi-variable systems
In the absence of...
Absolute Motion Analysis- General Plane Motion
As the drone's propellers rotate, an upward force is generated that counteracts the force of gravity, enabling the drone to lift off from the ground. This initial movement of the drone is along a straight path, representing a form of translational motion. In this phase, every point on the drone...
Planar Rigid-Body Motion
Planar motion is typically divided into three distinct categories. The first is rectilinear translation, demonstrated by a subway train that moves along...
