Related Experiment Video

Updated: Jun 10, 2026

Utilizing vmTracking to Improve the Accuracy of Multi-Animal Pose Estimation in Rodent Social Behavior Studies
07:34

Utilizing vmTracking to Improve the Accuracy of Multi-Animal Pose Estimation in Rodent Social Behavior Studies

Published on: November 7, 2025

Language Supervised Multi-Camera Multi-Object Tracking

Kaige Mao, Xiaopeng Hong, Xiaopeng Fan

    IEEE Transactions on Image Processing : a Publication of the IEEE Signal Processing Society
    |June 8, 2026
    PubMed

    Abstract:

    Recent multi-camera multi-object tracking (MCMOT) algorithms are primarily trained using per-detection identity annotations, which are complicated to obtain. In contrast, labeling a language description per-object is a more natural and human-friendly way. In this paper, we explore MCMOT in a language-supervised manner (LS-MCMOT) and propose a novel approach LaVST, which performs language-to-vision weakly-supervised learning based on reliable pseudo-labels generated via tracklet-level cross-modality matching. In addition, we design an ID-aware projection self-correction mechanism to correct inaccurate image-to-ground projection in a self-supervised manner. The models trained with our approach exhibit promising performance in LS-MCMOT. Surprisingly, they perform favorably against state-of-the-art identity-supervised methods, especially in cross-dataset evaluation (with an average gain by 20.0% in IDF1), underscoring the potential of language annotations in MCMOT. Codes and language annotations will be available here.

    More Related Videos

    A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
    12:39

    A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers

    Published on: January 18, 2020

    Related Experiment Videos

    Last Updated: Jun 10, 2026

    Utilizing vmTracking to Improve the Accuracy of Multi-Animal Pose Estimation in Rodent Social Behavior Studies
    07:34

    Utilizing vmTracking to Improve the Accuracy of Multi-Animal Pose Estimation in Rodent Social Behavior Studies

    Published on: November 7, 2025

    A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
    12:39

    A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers

    Published on: January 18, 2020

    Related Concept Videos

    Multi-input and Multi-variable systems01:22

    Multi-input and Multi-variable systems

    Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
    In the absence of...

    Articles linked to this work by shared authors, journal, and citation graph.

    Standardization and optimization of force platform testing for balance assessment in cervical spondylotic myelopathy.

    Frontiers in neuroscience·2026

    Precise resistance training alleviates osteosarcopenia via the GLP-1/myostatin path: a 32-week randomized controlled trial and molecular mechanism study.

    Archives of physiology and biochemistry·2026

    Robust Neural Depth Prediction From Uncalibrated Small Motion Clip.

    IEEE transactions on image processing : a publication of the IEEE Signal Processing Society·2026

    Near-infrared fluorescence imaging improves biliary visualization and perioperative outcomes in difficult laparoscopic cholecystectomy: implications for future clinical studies.

    Surgical endoscopy·2026

    Surgical video workflow analysis via visual-language learning.

    npj health systems·2026

    An ECM-mimetic hydrogel for disc repair: reconstituting hypoxia and alleviating NPC senescence to halt intervertebral disc degeneration.

    Journal of nanobiotechnology·2026

    Blind Attribute Quality Enhancement for G-PCC Compressed Dynamic Point Clouds.

    IEEE transactions on image processing : a publication of the IEEE Signal Processing Society·2026

    KN-LIO: Kinematics and Neural Field Coupled LiDAR-Inertial Odometry.

    IEEE transactions on image processing : a publication of the IEEE Signal Processing Society·2026

    Second-Order Visual Attention Prediction in Multiagent Videos via Emotion-Aware Belief Inference.

    IEEE transactions on image processing : a publication of the IEEE Signal Processing Society·2026

    DSES-Diff: A Dynamic-Spectral and Efficient-Spatial Diffusion-Based Foundation Model for Scalable Hyperspectral and Multispectral Image Fusion.

    IEEE transactions on image processing : a publication of the IEEE Signal Processing Society·2026

    DFO-CAM: Dual-Flow Optimization Towards Faithful Visual Explanations.

    IEEE transactions on image processing : a publication of the IEEE Signal Processing Society·2026

    Focus on the Hard: SAM-Based Prompt-Guided Refinement for RGB-T Semantic Segmentation.

    IEEE transactions on image processing : a publication of the IEEE Signal Processing Society·2026

    Artificial intelligence and vascular surgeons in patient communication: A comparative analysis of intelligibility and clinical appropriateness in acute deep vein thrombosis.

    Phlebology·2026

    Outcome of Unilateral Transverse Cordotomy or Kashima Operation for Bilateral Vocal Cord Immobility.

    Mymensingh medical journal : MMJ·2026

    Large Language Models as Diagnostic Tools in Nephrology: Helpful, With Limitations.

    Deutsches Arzteblatt international·2026

    'I think the website has completely reshaped our community': The history of the development of a trans-led online health resource.

    International journal of transgender health·2026

    Cloud-Based and Locally Deployed Language Models in Nursing and Health Care: An AI Act-Aligned Framework.

    JMIR medical informatics·2026

    Conditional validity in LLM-mediated L2 assessment: an argument-based systematic review and meta-analysis.

    Frontiers in research metrics and analytics·2026
    See all related articles
    JoVE
    x logofacebook logolinkedin logoyoutube logo
    ABOUT JoVE
    OverviewLeadershipBlogJoVE Help Center
    AUTHORS
    Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
    LIBRARIANS
    TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
    RESEARCH
    JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
    EDUCATION
    JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
    Terms & Conditions of Use
    Privacy Policy
    Policies
    Jove
    Visualize
    Contact Us