Related Experiment Video
Updated: Aug 20, 2025

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
Video Action Recognition Using Motion and Multi-View Excitation with Temporal Aggregation
Yuri Yudhaswana Joefrie1,2, Masaki Aono1
1Department of Computer Science and Engineering, Toyohashi University of Technology, 1-1 Tenpaku-cho, Toyohashi 441-8580, Japan.
This study introduces the META block for efficient video action recognition, improving spatial, temporal, and motion feature learning. The novel approach enhances performance on key datasets compared to existing methods.
Area of Science:
- Computer Vision
- Machine Learning
- Artificial Intelligence
Background:
- Video action recognition relies on effective spatiotemporal and motion feature representations.
- Existing methods like 3D CNNs are computationally expensive, while (1+2)D CNNs may neglect motion information.
Purpose of the Study:
- To propose a novel building block, the META block, for efficient and faithful capture of spatial, temporal, and motion features in videos.
- To address the limitations of existing 3D CNN and (1+2)D CNN approaches in video action recognition.
Main Methods:
- Introduction of the META block, comprising Motion Excitation (ME), Multi-view Excitation (MvE), and Densely Connected Temporal Aggregation (DCTA).
- ME encodes feature-level frame differences.
- MvE enriches spatiotemporal features using adaptive multi-view representations.
- DCTA models long-range temporal dependencies.
- Integration of the META block into a 2D ResNet-50 architecture.
Main Results:
- The proposed META block architecture demonstrates superior performance over previous CNN-based methods on the Something-Something v1 and Jester datasets, measured by 'Val Top-1 %'.
- Competitive results were achieved on the Moment-in-Time Mini dataset.
Conclusions:
- The META block offers a more faithful and efficient method for capturing spatial, temporal, and motion features in video action recognition.
- The proposed approach represents a significant advancement over existing CNN-based techniques for video understanding tasks.
Related Concept Videos
Relative Motion Analysis using Rotating Axes
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it...
Absolute Motion Analysis- General Plane Motion
As the drone's propellers rotate, an upward force is generated that counteracts the force of gravity, enabling the drone to lift off from the ground. This initial movement of the drone is along a straight path, representing a form of translational motion. In this phase, every point on the...
Relative Motion Analysis - Velocity
When an external force is exerted, it sets the crank into a rotational movement. This, in turn, instigates the motion of the connecting rod, leading to what is referred to as a general plane motion. This process involves two key points - point A on the connecting rod...
Relative Motion Analysis - Acceleration
Relative Motion Analysis using Rotating Axes - Acceleration
Time differentiation is...
Relative Motion Analysis using Rotating Axes-Problem Solving
Here, in order to determine the magnitude of velocity and acceleration for point...

