VIDHALLUC: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding

Chaoyu Li1, Eun Woo Im1, Pooyan Fazli1

  • 1Arizona State University.

Proceedings. IEEE Computer Society Conference on Computer Vision and Pattern Recognition
|September 5, 2025
PubMed
Summary

Multimodal large language models (MLLMs) struggle with video hallucinations. We introduce VIDHALLUC, a benchmark, and DINO-HEAL, a method improving MLLM accuracy in video understanding.