KE-VUM:知识增强的视频理解模型,用于细粒度的电影描述
Xiaojing Gu1, Gen Xu2, Xiaolu Zhang2
1University of Chinese Academy of Sciences, Beijing, China.
概括
本研究介绍了KE-VUM,这是一种用于详细描述电影片段的新型模型. 它利用脚本知识显著提高叙事连贯性和描述性准确性,优于现有方法.
科学领域:
- 人工智能的人工智能
- 计算机视觉 计算机视觉
- 自然语言处理自然语言处理.
背景情况:
- 精细的电影描述受到有限的注释和交叉模式的语义差距的阻碍,特别是在复杂的叙述中.
- 当前的视频描述模型往往无法利用结构化的脚本知识,导致肤浅的语义和字符错误识别.
研究的目的:
- 开发一个知识增强的视频理解模型 (KE-VUM),以改善细粒度的电影描述.
- 解决视频理解中的语义差距和叙事复杂性的挑战.
- 通过整合脚本知识,提高视频描述的时间和因果连贯性.
主要方法:
- KE-VUM 采用基于视频的多式联络骨干和脚本知识库之间的协作推理.
- 一个双阶段的脚本导向优化改进了草稿描述:通过知识图进行实体/关系校正,并与情节进展保持一致的叙事重建.
- 介绍了MovieClip对讲述丰富的剪辑和VD-Eval的基准标准,这是一个基于LLM的框架,用于评估语义准确性得分 (SAS) 和叙述一致性得分 (NCS).
主要成果:
- 在MovieClip基准测试中,KE-VUM表现优于强的基线.
- 在BERTScore F1中实现了6.7%的绝对改善,这表明描述性准确性得到了提高.
- 在延迟叙事场景中展示了强大的表现,提高了语义准确性和叙事连贯性.
结论:
- KE-VUM有效地整合了脚本知识,以获得更准确和更连贯的细粒度视频描述.
- 拟议的模型克服了现有方法的局限性,通过利用结构化知识来更深入地理解视频.
- 电影片段基准和VD-Eval为推进叙事视频理解研究提供了宝贵的资源.
相关概念视频
Stereotype Content Model
15.6K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
15.6K
Observational Learning
1.1K
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
1.1K
Relative Motion Analysis - Velocity
844
A stroke engine has a slider-crank mechanism that converts rotational motion from the crank into linear motion of the slider or vice versa. This mechanism consists of three main parts: the crank, the connecting rod, and the slider.
When an external force is exerted, it sets the crank into a rotational movement. This, in turn, instigates the motion of the connecting rod, leading to what is referred to as a general plane motion. This process involves two key points - point A on the connecting rod...
When an external force is exerted, it sets the crank into a rotational movement. This, in turn, instigates the motion of the connecting rod, leading to what is referred to as a general plane motion. This process involves two key points - point A on the connecting rod...
844
Induced-fit Model
90.3K
Most chemical reactions in cells require enzymes—biological catalysts that speed up the reaction without being consumed or permanently changed. They reduce the activation energy needed to convert the reactants into products. Enzymes are proteins, that usually work by binding to a substrate—a reactant molecule that they act upon.
Enzymes exhibit substrate specificity, meaning that they can only bind to certain substrates. This is mainly determined by the shape and chemical...
Enzymes exhibit substrate specificity, meaning that they can only bind to certain substrates. This is mainly determined by the shape and chemical...
90.3K
