KE-VUM:ファイングレイン映画記述のための知識強化ビデオ理解モデル
Xiaojing Gu1, Gen Xu2, Xiaolu Zhang2
1University of Chinese Academy of Sciences, Beijing, China.
まとめ
本研究では、詳細な映画クリップ記述のための新しいモデルであるKE-VUMを紹介します。スクリプト知識を活用して、物語の一貫性と記述精度を大幅に向上させ、既存の方法を上回っています。
科学分野:
- 人工知能;コンピュータビジョン;自然言語処理
背景:
- ファイングレイン映画記述は、アノテーションの制限と、特に複雑な物語におけるクロスモーダル意味ギャップによって妨げられています。現在のビデオ記述モデルは、構造化されたスクリプト知識を利用できず、表層的な意味とキャラクターの誤同定につながることがよくあります。
研究 の 目的:
- 改善されたファイングレイン映画記述のための知識強化ビデオ理解モデル(KE-VUM)を開発すること。ビデオ理解における意味のギャップと物語の複雑さの課題に対処すること。スクリプト知識を統合することにより、ビデオ記述における時間的および因果的整合性を強化すること。
主な方法:
- KE-VUMは、ビデオベースのマルチモーダルバックボーンとスクリプト知識ベース間の協調的推論を採用しています。2段階のスクリプトガイド最適化により、下書き記述が洗練されます。エンティティ/関係の修正(知識グラフ経由)とプロットの進行に合わせた物語の再構築。物語が豊富なクリップのためのMovieClipベンチマークと、意味精度スコア(SAS)と物語の一貫性スコア(NCS)を評価するためのLLMベースのフレームワークであるVD-Evalを導入しました。
主要な成果:
- KE-VUMは、強力なベースラインと比較して、MovieClipベンチマークで優れたパフォーマンスを示しました。BERTScore F1で6.7%の絶対的な改善を達成し、記述精度の向上が示されました。意味精度と物語の一貫性の両方を向上させる遅延物語シナリオで堅牢なパフォーマンスを示しました。
結論:
- KE-VUMは、スクリプト知識を効果的に統合し、より正確で一貫性のあるファイングレインビデオ記述を実現します。提案されたモデルは、構造化された知識を活用してビデオの理解を深めることにより、既存の方法の限界を克服します。MovieClipベンチマークとVD-Evalは、物語ビデオ理解の研究を進歩させるための貴重なリソースを提供します。
関連する概念動画
Stereotype Content Model
15.6K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
15.6K
Observational Learning
1.1K
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
1.1K
Relative Motion Analysis - Velocity
844
A stroke engine has a slider-crank mechanism that converts rotational motion from the crank into linear motion of the slider or vice versa. This mechanism consists of three main parts: the crank, the connecting rod, and the slider.
When an external force is exerted, it sets the crank into a rotational movement. This, in turn, instigates the motion of the connecting rod, leading to what is referred to as a general plane motion. This process involves two key points - point A on the connecting rod...
When an external force is exerted, it sets the crank into a rotational movement. This, in turn, instigates the motion of the connecting rod, leading to what is referred to as a general plane motion. This process involves two key points - point A on the connecting rod...
844
Induced-fit Model
90.3K
Most chemical reactions in cells require enzymes—biological catalysts that speed up the reaction without being consumed or permanently changed. They reduce the activation energy needed to convert the reactants into products. Enzymes are proteins, that usually work by binding to a substrate—a reactant molecule that they act upon.
Enzymes exhibit substrate specificity, meaning that they can only bind to certain substrates. This is mainly determined by the shape and chemical...
Enzymes exhibit substrate specificity, meaning that they can only bind to certain substrates. This is mainly determined by the shape and chemical...
90.3K
