皮尔辛眼:双空间视频暴力检测与高波视觉语言指导
IEEE transactions on pattern analysis and machine intelligence
|October 6, 2025
概括
皮尔辛眼通过结合欧几里德式和代式学习来增强视频暴力检测 (VVD). 这种新的方法可以更好地区分类似事件,并使用先进的人工智能准确识别模两可的暴力.
科学领域:
- 计算机科学 计算机科学
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 目前监督较弱的视频暴力检测 (VVD) 方法因层次模型和训练数据的局限性而与视觉上类似的事件作斗争.
- 在VVD中常见的欧几里德表示学习,缺乏在语义上不同但视觉上相似的事件之间进行细微歧视的能力.
研究的目的:
- 介绍PiercingEye,一个双空间学习框架,集成欧几里德几何和形几何,在VVD中提供优越的歧视性特征表示.
- 通过协同不同几何空间来增强事件层次结构和特征交互的建模.
主要方法:
- 开发了一种新的双空间学习框架 (PiercingEye),结合了欧几里德几何和过度几何.
- 实施了一种层敏感的超标聚合策略,用于层次事件建模的超标迪里克莱特能量约束.
- 引入了功能交互的跨空间注意力机制,并利用了大型语言模型来生成模两可的事件描述,训练了超标视觉语言对比损失.
主要成果:
- 在XD-Violence和UCF-Crime基准指标上,PiercingEye取得了最先进的表现.
- 展示了优越的细粒度暴力检测能力,特别是在一个精心策划的模两可的事件子集上.
- 提出的方法有效地解决了现有的VVD方法在处理模两可和视觉上相似的事件方面的局限性.
结论:
- 双空间学习框架PiercingEye通过改进特征表示和处理模两可的样本,显著提升了视频暴力检测.
- 超标几何学和视觉语言建模的整合为更强大,更准确的视频分析提供了有希望的方向.
- 在具有挑战性的数据集上PiercingEye的成功凸显了其在需要精确事件检测的现实应用中的潜力.
相关概念视频
Depth Perception and Spatial Vision
1.8K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
1.8K
Collisions in Multiple Dimensions: Problem Solving
5.3K
In multiple dimensions, the conservation of momentum applies in each direction independently. Hence, to solve collisions in multiple dimensions, we should write down the momentum conservation in each direction separately. To help understand collisions in multiple dimensions, consider an example.
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
5.3K
Collisions in Multiple Dimensions: Introduction
6.5K
It is far more common for collisions to occur in two dimensions; that is, the initial velocity vectors are neither parallel nor antiparallel to each other. Let's see what complications arise from this. The first idea is that momentum is a vector. Like all vectors, it can be expressed as a sum of perpendicular components (usually, though not always, an x-component and a y-component, and a z-component if necessary). Thus, when the statement of conservation of momentum is written for a...
6.5K
Relative Motion Analysis using Rotating Axes-Problem Solving
702
Consider a crane whose telescopic boom rotates with an angular velocity of 0.04 rad/s and angular acceleration of 0.02 rad/s2. Along with the rotation, the boom also extends linearly with a uniform speed of 5 m/s. The extension of the boom is measured at point D, which is measured with respect to the fixed point C on the other end of the boom. For the given instant, the distance between points C and D is 60 meters.
Here, in order to determine the magnitude of velocity and acceleration for point...
Here, in order to determine the magnitude of velocity and acceleration for point...
702
Vision
59.3K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
59.3K
Relative Motion Analysis using Rotating Axes
879
Consider a component AB undergoing a linear motion. Along with a linear motion, point B also rotates around point A. To comprehend this complex movement, position vectors for both points A and B are established using a stationary reference frame.
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it...
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it...
879


