VIDHALLUC:ビデオ理解のためのマルチモダルの大型言語モデルにおける時間幻覚の評価
Chaoyu Li1, Eun Woo Im1, Pooyan Fazli1
1Arizona State University.
まとめ
マルチモダルの大きな言語モデル (MLLMs) は ビデオ幻覚と闘っています ビデオ理解におけるMLLMの精度を向上させる方法であるDINO-HEALを紹介しています.
科学分野:
- 人工知能
- コンピュータ・ビジョン
- 機械学習
背景:
- マルチモダル大型言語モデル (MLLM) は,高度なビデオ理解能力を実証しています.
- MLLMのビデオ分析において,不正確なコンテンツの生成という幻覚は重要な課題です.
- 既存のMLLMのビジュアルエンコーダーは,視覚的に似ているが,意味学的に異なるビデオコンテンツを混同することが多い.
研究 の 目的:
- ビデオ理解における MLLM 幻覚の評価のための 最大の基準である VIDHALLUC を導入します
- 幻覚に対するMLLMの脆弱性を評価する
- DINO-HEALを提案する ビデオの幻覚を緩和する 新しい方法
主な方法:
- VIDHALLUCのベンチマークは5,002本のペア化されたビデオで構成されています.
- VIDHALLUCの基準で様々なMLLMの総合的なテストを行いました.
- DINO-HEALの導入,機能の重み付けのためにDINOv2の空間的突出性を使用するトレーニングフリーテクニック.
主要な成果:
- ほとんどのMLLMは,アクション,タイムシーケンス,シーンのトランジションの次元にわたって有意な幻覚を示します.
- DINO- HEALはMLLMの幻覚を効果的に軽減する.
- DINO- HEALは,タスクの間で幻覚を軽減する平均3. 02%の改善を達成しました.
結論:
- MLLMはビデオ理解作業で 幻覚に敏感です
- VIDHALLUCの基準は,これらの限界を評価し,対処するための重要なリソースを提供します.
- DINO-HEALは,再訓練なしで,MLLMのビデオ幻覚を減らすための実用的で効果的な解決策を提供します.
関連する概念動画
Higher Mental Functions of the Brain: Language
1.0K
Language is a system of communication that allows the expression of thoughts, ideas, and feelings. The brain processes language in both hemispheres.
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
1.0K
Chunking and Rehearsal in Sensory Memory
293
Improving short-term memory can be achieved through techniques like chunking and rehearsal. Chunking involves organizing information into larger, more manageable units. This technique is particularly useful for information that exceeds the typical memory span of between five and nine items. For instance, logging into an online account with a password like "ta89vq0179gz" involves grouping letters and numbers into three chunks—ta89, vq01, and 79gz. It makes large amounts of...
293


