Related Experiment Video
Updated: Jul 20, 2025

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
9.0K
Unsupervised Video Anomaly Detection Based on Similarity with Predefined Text Descriptions
Jaehyun Kim1, Seongwook Yoon1, Taehyeon Choi1
1School of Electrical Engineering, Korea University, Seoul 02841, Republic of Korea.
Sensors (Basel, Switzerland)
|July 29, 2023
Summary
This study introduces a novel unsupervised video anomaly detection method using text descriptions and the CLIP model. It achieves strong performance, outperforming existing unsupervised techniques without requiring extensive dataset labeling.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Traditional video anomaly detection relies heavily on video data.
- Real-world applications benefit from incorporating human domain knowledge, often expressed as text.
- Unlabeled video datasets are abundant, presenting an opportunity for unsupervised learning.
Purpose of the Study:
- To explore the use of text descriptions for unsupervised video anomaly detection.
- To develop a method that leverages text to identify abnormal situations in videos without prior labeling.
- To improve anomaly detection performance by integrating textual domain knowledge with visual data.
Main Methods:
- Utilized large language models to generate text descriptions for video content.
- Employed the CLIP (Contrastive Language-Image Pre-training) visual language model to compute cosine similarity between video frames and text descriptions.
- Introduced a text-conditional similarity refinement using an unlabeled dataset and a triplet loss function.
Main Results:
- The proposed method significantly outperforms existing unsupervised methods on the ShanghaiTech and UCFcrime datasets, achieving 8% and 13% higher AUC scores, respectively.
- Demonstrated comparable performance to weakly supervised methods in detecting abnormal videos, despite not requiring manual labeling.
- Achieved 17% and 5% better AUC scores on abnormal videos compared to weakly supervised approaches, highlighting its effectiveness.
Conclusions:
- Text descriptions can be effectively integrated into unsupervised video anomaly detection frameworks.
- The proposed method offers a computationally efficient alternative to traditional approaches, avoiding optical flow or multi-frame analysis.
- This research validates the potential of using readily available text descriptions to enhance unsupervised anomaly detection in video data.
Keywords:
CLIPabnormal videoembedding spacefine-tuning of pre-trained modelslarge language modelslarge vision and language modelssimilarity measuretext descriptionsunsupervised video anomaly detectionMore Related Videos
Related Concept Videos
Unusual Results
3.2K
Unusual results are those that have a very low chance of occurring. Unusual results can be identified using probabilities and the range rule of thumb. In problems involving probability, unusual results can be observed in 2 instances – an unusually high number of successes or an unusually low number of successes.
According to the range rule of thumb, any value above or below two standard deviations, 2σ from the mean, μ is considered unusual.
Maximum unusual value =...
According to the range rule of thumb, any value above or below two standard deviations, 2σ from the mean, μ is considered unusual.
Maximum unusual value =...
3.2K
Difference from Background: Limit of Detection
6.4K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
6.4K
Detection of Black Holes
2.2K
Although black holes were theoretically postulated in the 1920s, they remained outside the domain of observational astronomy until the 1970s.
Their closest cousins are neutron stars, which are composed almost entirely of neutrons packed against each other, making them extremely dense. A neutron star has the same mass as the Sun but its diameter is only a few kilometers. Therefore, the escape velocity from their surface is close to the speed of light.
Not until the 1960s, when the first neutron...
Their closest cousins are neutron stars, which are composed almost entirely of neutrons packed against each other, making them extremely dense. A neutron star has the same mass as the Sun but its diameter is only a few kilometers. Therefore, the escape velocity from their surface is close to the speed of light.
Not until the 1960s, when the first neutron...
2.2K
Classification of Signals
532
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
532

