Related Experiment Video
Updated: Feb 20, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
VL-HTR: Learning Human-Target Representation From Vision-Language Model
None:
Human-gaze-target prediction aims to predict the target point or object that humans are looking at in images. However, existing methods predominantly rely on vision-only features, which often struggle to capture the semantic context of small or occluded objects and lack explicit priors for precise head direction regression, leading to slow convergence and suboptimal performance. Therefore, we introduce VL-HTR, a novel vision-language learning method for human-target representation, which integrates multimodal knowledge from vision-language models (VLMs) to construct robust human-target relationships. Unlike traditional approaches, extracting multimodal features via pretrained VLMs enhances the model's grasp of human-target knowledge through the learnable target class and direction context. Then, a language-guided query alignment (LQA) module is introduced to improve the semantic-aware object representation capability through vision-language query alignment. Finally, to accelerate the gaze point regression learning process, we design a language-guided direction prediction (LDP) module to introduce multimodal human gaze direction priors, thereby facilitating the human-target relationship construction. Extensive validations across two distinct tasks, i.e., gaze object prediction (GOP) and gaze target estimation, involving five challenging benchmarks, demonstrating that VL-HTR achieves superior performance and much faster training convergence.
Related Concept Videos
Vision
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Language and Cognition