Related Experiment Video
Updated: Sep 19, 2025

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
Recognizing human-object interactions in videos with the supervision of natural language
Qiyue Li1, Xuemei Xie2, Jin Zhang2
1School of Communications and Information Engineering, Xi'an University of Posts and Telecommunications, Xi'an, Shaanxi 710121, PR China.
Abstract:
Existing models for recognizing human-object interaction (HOI) in videos mainly rely on visual information for reasoning and generally treat recognition tasks as traditional multi-classification problems, where labels are represented by numbers. This supervised learning method discards semantic information in the labels and ignores advanced semantic relationships between actual categories. In fact, natural language contains a wealth of linguistic knowledge that humans have distilled about human-object interaction, and the category text contains a large amount of semantic relationships between texts. Therefore, this paper introduces human-object interaction category text features as labels and proposes a natural language supervised learning model for human-object interaction by using natural language to supervise visual feature learning to enhance visual feature expression capability. The model applies contrastive learning paradigm to human-object interaction recognition, using an image-text paired pre-training model to obtain individual image features and interaction category text features, and then using a spatial-temporal mixed module to obtain high semantic combination-based human-object interaction spatial-temporal features. Finally, the obtained visual interaction features and category text features are compared for similarity to infer the correct video human-object interaction category. The model aims to explore the semantic information in human-object interaction category label text and use a large number of image-text paired samples trained by a multi-modal pre-training model to obtain visual and textual correspondence to enhance the ability of video human-object interaction recognition. Experimental results on two human-object interaction datasets demonstrate that our method achieves the state-of-the-art performance, e.g., 93.6% and 93.1% F1 Score for Sub-activity and Affordance on CAD-120 dataset.
Related Concept Videos
Automatic Processing and Automatic Social Behavior
Observational Learning
Natural and Artificial Concepts

