Related Experiment Video
Updated: Jul 18, 2026

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
TrCLIP-VAD : Weak supervised video anomaly detection by improving CLIP training with text rewriting
Shengjie Shen1, Ziteng Guo1, Yahui Li1
1School of Computer Science and Technology, Xinjiang University, Urumqi, China.
Abstract:
The existing video anomaly detection (VAD) methods based on the contrastive language-image pre-training (CLIP) framework primarily rely on spatial-temporal features and label prompts, often overlooking high-level semantic information, such as text descriptions of abnormal behavior. Moreover, while data augmentation techniques are commonly used to expand video frames, text inputs typically remain unchanged throughout the training process, limiting the exposure of different texts to the same image. In this paper, we propose a text rewriting based CLIP model called TrCLIP-VAD. It first generates source caption descriptions for each video and then utilizes the In-Context Learning (ICL) capabilities of large language models (LLMs) to rewrite these captions, thereby enhancing semantic expression. Subsequently, we randomly select a caption and fuse it with visual features to detect anomalies. In addition, we introduce the Local-Global Multi-scale Mamba (LGM-Mamba) module. The LGM-Mamba module aims to capture the local-global temporal dependence of frame sequence and promote the deep fusion of visual and caption features to improve the detection performance. The experimental results show that TrCLIP-VAD achieves state-of-the-art performance on two widely used datasets, XD-Violence and UCF-Crime. Specifically, TrCLIP-VAD achieves 86.05% AP scores on XD-Violence dataset and 88.59% AUC scores on UCF-Crime dataset. The code is released at https://github.com/vpsg-research/TrCLIP-VAD.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy