Related Experiment Video
Updated: Feb 8, 2026

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
A lightweight convolutional neural network architecture for violence detection in video sequences
Bhawana Tyagi1, Richa Jain2, Pankaj Jain3
1School of Computer Science and Engineering, VIT University, Vellore, Tamil Nadu, India. bhawana1988@gmail.com.
Abstract:
The escalation of violent incidents in high-density public environments such as political assemblies, concerts, and sports arenas necessitates the development of computationally efficient and accurate real-time violence detection frameworks. Prompt identification of aggressive events from continuous surveillance video streams is critical for initiating rapid countermeasures. However, the task is inherently complex due to spatiotemporal scene variations, illumination inconsistencies, and the intensive computational cost of processing high-dimensional video data. This study introduces a lightweight deep convolutional neural network (CNN) architecture derived from MobileNetV2, optimized through depthwise separable convolutions and inverted residual bottlenecks to achieve significant parameter reduction without compromising classification efficacy. The proposed framework processes video streams by extracting and preprocessing frames (224 × 224 resolution, normalization, augmentation) to enhance generalization and mitigate overfitting. The model was trained and evaluated on two benchmark datasets: the Real-Life Violence Situations Dataset (RLVSD) and the Hockey Fight Dataset (HFD), encompassing balanced classes of violent and non-violent sequences. Empirical evaluation indicates superior performance, attaining 97% accuracy on RLVSD and 94% on HFD, with corresponding gains in precision, recall, and F1-score compared to conventional CNN architectures. Computational profiling confirms substantial efficiency improvements, enabling inference at real-time frame rates on resource-constrained hardware. The proposed methodology demonstrates that optimized lightweight architectures can deliver high-accuracy violence detection while significantly reducing computational overhead. These characteristics make the approach highly deployable in real-world surveillance systems. Future research will focus on temporal feature integration via 3D CNNs or transformer-based models and cross-domain adaptability to heterogeneous video sources.
Related Concept Videos
Sequence Networks of Rotating Machines
Zero-sequence current induces a voltage drop across the generator's neutral impedance and other...
Convolution Properties II
The width property indicates that if the durations of input signals are T1 and T2, then the width of the output response equals the sum of both durations, irrespective of the shapes of the two functions. For instance, convolving two rectangular pulses with durations of 2 seconds and 1 second results in a function with a width of 3 seconds.
The area property asserts that the area under the...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Convolution Properties I
The commutative property reveals that the input and the impulse response of an LTI (Linear Time-Invariant) system can be interchanged without affecting the output:
Network Covalent Solids
To break or to melt a covalent network solid, covalent bonds must be broken. Because covalent bonds are relatively strong, covalent network solids are typically...
Neural Regulation

