Related Experiment Video
Updated: Feb 12, 2026

Defining the Role Of Language in Infants' Object Categorization with Eye-tracking Paradigms
Published on: February 8, 2019
Monocular Multi-Object 3D Visual Language Tracking
This study introduces the first method for multi-object 3D visual language tracking (VLT) using monocular video. The new approach enables precise 3D tracking of multiple objects from single camera feeds, overcoming previous limitations.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Robotics
Background:
- Existing Visual Language Tracking (VLT) methods are limited to 2D or single-object 3D tracking.
- Current 3D multi-object tracking relies on sensor data lacking language descriptions.
- Redundant language descriptions hinder precise multi-object localization in VLT.
Purpose of the Study:
- To extend Visual Language Tracking (VLT) to multi-object 3D tracking using monocular video.
- To introduce a novel framework, dataset, and neural model for this task.
- To enable robust 3D tracking of multiple objects from monocular video using language guidance.
Main Methods:
- Introduced the Monocular Multi-object 3D Visual Language Tracking (MoMo-3DVLT) task and dataset (MoMo-3DRoVLT).
- Developed MoMo-3DVLTracker, a neural model with multimodal feature extraction and a language encoder-decoder.
- Integrated a differentiable linked-memory mechanism with depth-guided, language-conditioned reasoning.
Main Results:
- The proposed MoMo-3DVLTracker achieves state-of-the-art performance on the MoMo-3DRoVLT dataset.
- Demonstrated superior multi-object 3D tracking accuracy using monocular video compared to existing methods.
- The custom dataset provides extensive 3D annotations and natural language descriptions for VLT research.
Conclusions:
- This work presents the first successful extension of VLT to multi-object 3D tracking with monocular video.
- The developed framework, dataset, and model establish a new baseline for this challenging task.
- The approach offers a promising direction for real-world applications requiring language-guided 3D object tracking.
More Related Videos
07:36Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
06:07Exploring Infant Sensitivity to Visual Language using Eye Tracking and the Preferential Looking Paradigm
Published on: May 15, 2019
Related Concept Videos
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Velocity of an Object
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...