VIDHALLUC:在视频理解的多模式大语言模型中评估时间幻觉
Chaoyu Li1, Eun Woo Im1, Pooyan Fazli1
1Arizona State University.
概括
多模式大语言模型 (MLLMs) 在视频幻觉中扎. 我们介绍了VIDHALLUC,一个基准,以及DINO-HEAL,一种提高视频理解MLLM准确性的方法.
科学领域:
- 人工智能
- 计算机视觉
- 机器学习
背景情况:
- 多模式大语言模型 (MLLM) 展示了先进的视频理解能力.
- 幻觉,不准确的内容的产生,是MLLM视频分析的一个重大挑战.
- 现有的MLLM视觉编码器经常混视觉上相似但语义上不同的视频内容.
研究的目的:
- 介绍VIDHALLUC,这是评估MLLM视频理解中的最大基准.
- 通过行动,时间序列和场景过渡来评估MLLM对幻觉的脆弱性.
- 提出DINO-HEAL, 这是一种减轻视频幻觉的新方法.
主要方法:
- 开发了包括5002个配对视频的VIDHALLUC基准.
- 在VIDHALLUC基准上对各种MLLM进行全面测试.
- 介绍DINO-HEAL,一种使用DINOv2空间突出度进行特征重量化的无训练技术.
主要成果:
- 大多数MLLM在动作,时间序列和场景过渡维度上都表现出显著的幻觉.
- DINO-HEAL有效地减少了MLLM患者的幻觉.
- 在不同任务中,DINO- HEAL 平均改善了3. 02%的幻觉.
结论:
- 在视频理解任务中容易产生幻觉.
- 维达卢克基准为评估和解决这些局限性提供了关键资源.
- DINO-HEAL提供了一种实际有效的解决方案,
相关概念视频
Higher Mental Functions of the Brain: Language
1.0K
Language is a system of communication that allows the expression of thoughts, ideas, and feelings. The brain processes language in both hemispheres.
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
1.0K
Chunking and Rehearsal in Sensory Memory
293
Improving short-term memory can be achieved through techniques like chunking and rehearsal. Chunking involves organizing information into larger, more manageable units. This technique is particularly useful for information that exceeds the typical memory span of between five and nine items. For instance, logging into an online account with a password like "ta89vq0179gz" involves grouping letters and numbers into three chunks—ta89, vq01, and 79gz. It makes large amounts of...
293


