Related Experiment Video
Updated: Jan 10, 2026

08:32
Tracking Rats in Operant Conditioning Chambers Using a Versatile Homemade Video Camera and DeepLabCut
Published on: June 15, 2020
13.3K
Benchmarking Compact VLMs for Clip-Level Surveillance Anomaly Detection Under Weak Supervision
Kirill Borodin1, Kirill Kondrashov1, Nikita Vasiliev1
1Faculty of Information Technology, Moscow Technical University of Communication and Informatics, Moscow 111024, Russia.
Journal of Imaging
|November 26, 2025
Summary
Compact vision-language models (VLMs) offer a practical solution for CCTV anomaly detection, balancing accuracy and speed. Parameter-efficient fine-tuning enhances their reliability and consistency for real-time safety monitoring.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- CCTV safety monitoring requires anomaly detection with high accuracy and low latency.
- Weak supervision presents challenges for traditional anomaly detection methods.
- Compact vision-language models (VLMs) are explored for their potential in this domain.
Purpose of the Study:
- To investigate the effectiveness of parameter-efficiently adapted compact VLMs for CCTV anomaly detection.
- To establish a unified evaluation protocol for comparing different VLM approaches and baselines.
- To assess the trade-off between detection accuracy and per-clip latency.
Main Methods:
- A unified evaluation protocol was developed, standardizing preprocessing, prompting, dataset splits, metrics, and runtime settings.
- Compact VLMs were adapted using parameter-efficient fine-tuning.
- Performance was compared against training-free VLM pipelines and weakly supervised baselines.
- Metrics included accuracy, precision, recall, F1, ROC-AUC, and average per-clip latency.
Main Results:
- Parameter-efficiently adapted compact VLMs achieved performance comparable to or exceeding established methods.
- These models maintained competitive per-clip latency, crucial for real-time monitoring.
- Adaptation reduced prompt sensitivity, leading to more consistent behavior.
- A favorable accuracy-efficiency trade-off was demonstrated.
Conclusions:
- Parameter-efficient fine-tuning enables compact VLMs to function as reliable clip-level anomaly detectors.
- Compact VLMs offer a practical and efficient solution for CCTV safety monitoring under weak supervision.
- The proposed evaluation protocol ensures transparency and consistency in assessing anomaly detection methods.
Related Concept Videos
Difference from Background: Limit of Detection
8.0K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
8.0K
Detection of Gross Error: The Q Test
6.8K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.8K