Related Experiment Video
Updated: Jul 13, 2026

Video Movement Analysis Using Smartphones ViMAS: A Pilot Study
Published on: March 14, 2017
SSIM over MSE: A new perspective for video anomaly detection
Jin Fan1, Miao Chen2, Zhangyu Gu2
1Department of Computer Science and Technology, Hangzhou Dianzi University, Hangzhou, 310018, Zhejiang, China; Zhejiang Provincial Key Laboratory of Internet in Discrete Industries, Hangzhou Dianzi University, Hangzhou, 310018, Zhejiang, China; Research and Development Center of Transport Industry of New Generation of Artificial Intelligence Technology, Hangzhou, 310018, Zhejiang, China.
This study enhances video anomaly detection by aligning models with human perception using Structural Similarity Index (SSIM). This improves accuracy and interpretability in public safety applications.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Pattern Recognition
Background:
- Current video anomaly detection models often fail to align with human perception due to reliance on Mean Squared Error (MSE).
- These models prioritize texture over shape, leading to poor interpretability and performance limitations.
- Existing methods struggle to accurately identify anomalies as perceived by the Human Visual System (HVS).
Purpose of the Study:
- To optimize video anomaly detection models by incorporating human visual relevance.
- To improve the alignment between model-detected anomalies and human perception.
- To enhance model interpretability and overall performance in detecting abnormal video patterns.
Main Methods:
- Introduction of a novel Structural Similarity Index (SSIM) based loss function.
- Development of a new anomaly score calculation method utilizing SSIM.
- Integration of a spatial-temporal enhancement block in 3D convolution (STE-3D) for improved feature capture.
Main Results:
- The SSIM-based loss emphasizes shape information over texture, enhancing model interpretability.
- The SSIM-based anomaly score aligns better with human visual perception.
- The STE-3D block effectively captures spatial-temporal features and compensates for SSIM loss limitations.
- Experimental validation on benchmarks like UCSD Ped1/Ped2, CUHK Avenue, and ShanghaiTech demonstrates significant performance improvements.
Conclusions:
- The proposed approach, integrating SSIM loss and STE-3D, effectively improves video anomaly detection.
- Aligning models with human visual perception leads to more interpretable and accurate anomaly detection.
- The methods are lightweight and seamlessly integrate with existing 3D convolutional models.

