Related Experiment Video
Updated: Mar 2, 2026

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
9.7K
KE-VUM: Knowledge-enhanced video understanding model for fine-grained movie description
Xiaojing Gu1, Gen Xu2, Xiaolu Zhang2
1University of Chinese Academy of Sciences, Beijing, China.
Summary
This study introduces KE-VUM, a novel model for detailed movie clip descriptions. It leverages script knowledge to significantly improve narrative coherence and descriptive accuracy, outperforming existing methods.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Natural Language Processing
Background:
- Fine-grained movie description is hindered by limited annotations and cross-modal semantic gaps, particularly in complex narratives.
- Current video description models often fail to utilize structured script knowledge, leading to superficial semantics and character misidentification.
Purpose of the Study:
- To develop a knowledge-enhanced video understanding model (KE-VUM) for improved fine-grained movie description.
- To address challenges in semantic gaps and narrative complexity in video understanding.
- To enhance temporal and causal coherence in video descriptions by integrating script knowledge.
Main Methods:
- KE-VUM employs collaborative reasoning between a video-based multimodal backbone and a script knowledge base.
- A two-stage script-guided optimization refines draft descriptions: entity/relationship correction via a knowledge graph and narrative reconstruction aligned with plot progressions.
- Introduced the MovieClip benchmark for narrative-rich clips and VD-Eval, an LLM-based framework for evaluating Semantic Accuracy Score (SAS) and Narrative Coherence Score (NCS).
Main Results:
- KE-VUM demonstrated superior performance over strong baselines on the MovieClip benchmark.
- Achieved a 6.7% absolute improvement in BERTScore F1, indicating enhanced descriptive accuracy.
- Showcased robust performance in delayed-narrative scenarios, improving both semantic accuracy and narrative coherence.
Conclusions:
- KE-VUM effectively integrates script knowledge for more accurate and coherent fine-grained video descriptions.
- The proposed model overcomes limitations of existing methods by leveraging structured knowledge for deeper video understanding.
- The MovieClip benchmark and VD-Eval provide valuable resources for advancing research in narrative video understanding.
Related Concept Videos
Stereotype Content Model
15.6K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
15.6K
Observational Learning
1.1K
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
1.1K
Relative Motion Analysis - Velocity
844
A stroke engine has a slider-crank mechanism that converts rotational motion from the crank into linear motion of the slider or vice versa. This mechanism consists of three main parts: the crank, the connecting rod, and the slider.
When an external force is exerted, it sets the crank into a rotational movement. This, in turn, instigates the motion of the connecting rod, leading to what is referred to as a general plane motion. This process involves two key points - point A on the connecting rod...
When an external force is exerted, it sets the crank into a rotational movement. This, in turn, instigates the motion of the connecting rod, leading to what is referred to as a general plane motion. This process involves two key points - point A on the connecting rod...
844
Induced-fit Model
90.3K
Most chemical reactions in cells require enzymes—biological catalysts that speed up the reaction without being consumed or permanently changed. They reduce the activation energy needed to convert the reactants into products. Enzymes are proteins, that usually work by binding to a substrate—a reactant molecule that they act upon.
Enzymes exhibit substrate specificity, meaning that they can only bind to certain substrates. This is mainly determined by the shape and chemical...
Enzymes exhibit substrate specificity, meaning that they can only bind to certain substrates. This is mainly determined by the shape and chemical...
90.3K