语义意识的视频表示,用于识别少量拍摄的动作
Yutao Tang1, Benjamín Béjar2, René Vidal3
1Johns Hopkins University.
概括
这项研究介绍了语义意识的少数拍摄动作识别 (SAFSAR) 模型,该模型使用3D特征和文本来更好地识别动作. 萨夫萨尔简化了时间建模和特征融合,以在少数拍摄场景中提高性能.
科学领域:
- 计算机视觉 计算机视觉
- 机器学习 机器学习
- 人工智能的人工智能
背景情况:
- 短拍动作识别方法通常使用2D功能,并与时间建模和文本集成作斗争.
- 现有的方法需要复杂的组件和距离函数,这限制了它们在融合多式联运信息方面的有效性.
研究的目的:
- 提出一个语义意识的少射击动作识别 (SAFSAR) 模型,克服当前方法的局限性.
- 通过使用简化方法,在几次射击动作识别中实现最先进的性能.
主要方法:
- 利用3D特征提取器和一个有效的特征融合方案,用于分类.
- 引入一种新的方案,将文本语义编码为适应融合的视频表示.
- 鼓励视觉编码器通过集成的文本和视频功能对齐来提取语义上一致的特征.
主要成果:
- 与现有方法相比,SAFSAR模型显示出更高的性能.
- 在五个具有挑战性的少量行动认可基准上取得了显著的改进.
- 验证了使用3D功能和简化融合方案的有效性.
结论:
- 拟议的SAFSAR模型提供了一种简单而有效的解决方案,用于识别少数射击行动.
- 直接使用3D功能与自适应融合优于复杂的方法.
- 在SAFSAR中,可以对文字和视频功能进行紧的对齐和融合,以提高性能.
相关概念视频
Fixed Action Patterns
15.9K
A fixed action pattern (FAP) is a specific, hard-wired sequence of behaviors that occurs in response to an external stimulus, called a sign stimulus. The behavior is “fixed” because it is essentially unchangeable—proceeding similarly across individuals of a species every time it occurs.
15.9K
Force Classification
1.2K
Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
1.2K
Observational Learning
154
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
154
Stereotype Content Model
14.0K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.0K
State Space Representation
178
The frequency-domain technique, commonly used in analyzing and designing feedback control systems, is effective for linear, time-invariant systems. However, it falls short when dealing with nonlinear, time-varying, and multiple-input multiple-output systems. The time-domain or state-space approach addresses these limitations by utilizing state variables to construct simultaneous, first-order differential equations, known as state equations, for an nth-order system.
Consider an RLC circuit, a...
Consider an RLC circuit, a...
178
Chunking and Rehearsal in Sensory Memory
186
Improving short-term memory can be achieved through techniques like chunking and rehearsal. Chunking involves organizing information into larger, more manageable units. This technique is particularly useful for information that exceeds the typical memory span of between five and nine items. For instance, logging into an online account with a password like "ta89vq0179gz" involves grouping letters and numbers into three chunks—ta89, vq01, and 79gz. It makes large amounts of...
186


