Related Experiment Videos
PoinCLIP-VAD: a hyperbolic cross-modal fusion framework for video anomaly detection
Debi Prasad Senapati1, Santosh Kumar Pani1, Santos Kumar Baliarsingh2
1School of Computer Engineering, KIIT Deemed to be University, Campus 25, Chandaka Industrial Estate Patia, Bhubaneswar, Odisha, 751024, India.
None:
Weakly supervised video anomaly detection (WSVAD) is fundamentally constrained by the absence of frame-level annotations, which leads to noisy instance selection in Multiple Instance Learning (MIL) and weak correspondence between temporal video segments and semantic descriptions. Vision-language models address this by enabling cross-modal alignment between visual features and textual labels, but when these representations are learned in Euclidean space, they struggle to capture subtle semantic variations and often produce ambiguous instance ranking under weak supervision. To address this limitation, we propose PoinCLIP-VAD, a vision-language framework that performs cross-modal fusion in hyperbolic space. The model embeds visual and textual features into a shared Poincaré ball geometry, where non-linear distance scaling provides a more expressive representation of latent semantic relationships induced by cross-modal interactions, without relying on predefined hierarchical structures. This geometry-consistent formulation enables more reliable similarity estimation and better preserves distinctions between normal and anomalous patterns. The framework adopts a dual-block architecture consisting of a classification block for coarse anomaly scoring and a video-text alignment block for fine-grained correspondence using negative Poincaré distance. Extensive experiments on benchmark datasets demonstrate that PoinCLIP-VAD achieves an AUC of 90.62% on UCF-Crime and an AP of 86.93% on XD-Violence, confirming improved anomaly discrimination and more consistent cross-modal alignment under weak supervision.
Related Concept Videos
Absolute Motion Analysis- General Plane Motion
As the drone's propellers rotate, an upward force is generated that counteracts the force of gravity, enabling the drone to lift off from the ground. This initial movement of the drone is along a straight path, representing a form of translational motion. In this phase, every point on the drone...
Relative Motion Analysis using Rotating Axes
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it instrumental in...
Vector or Cross Product
Consider the cross product of two vectors. Imagine rotating the first vector about...
Cross Product
The magnitude of the cross product is obtained by multiplying the magnitude of both the vectors and the sine of the angle between them. This means that a larger angle between the vectors will lead to a greater magnitude of the cross product.
Cross Product and Its Geometry
Relative Motion Analysis using Rotating Axes-Problem Solving
Here, in order to determine the magnitude of velocity and acceleration for point...