Related Experiment Video
Updated: May 5, 2026

06:37
Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
4.5K
A generalized pyramid matching kernel for human action recognition in realistic videos
Jun Zhu1, Quan Zhou, Weijia Zou
1Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University, Shanghai 200240, China. zhangwenjun@sjtu.edu.cn.
Sensors (Basel, Switzerland)
|November 29, 2013
Summary
We introduce a generalized pyramid matching kernel (GPMK) to improve human action recognition in videos. This method effectively handles variations in realistic videos, outperforming existing techniques.
Area of Science:
- Computer Vision
- Machine Learning
- Artificial Intelligence
Background:
- Human action recognition is crucial for video analysis but faces challenges due to unconstrained sensing conditions.
- Realistic videos exhibit significant intra-class variations and inter-class ambiguities, hindering current vision-based systems.
- Existing methods struggle with the complexity and variability of real-world video data.
Purpose of the Study:
- To propose a novel Generalized Pyramid Matching Kernel (GPMK) for enhanced human action recognition.
- To address the limitations of traditional methods in handling diverse visual cues and spatial-temporal information.
- To develop a robust similarity metric for kernel-based classification of video clips.
Main Methods:
- A multi-channel "bag of words" representation is constructed from local spatial-temporal features.
- The GPMK extends the spatial-temporal pyramid matching (STPM) kernel by leveraging heterogeneous visual cues.
- Adaptive channel weights are computed using kernel target alignment, integrating prior and data-driven information.
Main Results:
- The GPMK demonstrates superior performance compared to the traditional STPM kernel on challenging datasets (Hollywood2, Youtube, HMDB51).
- The proposed method effectively handles intra-class variations and inter-class ambiguities in realistic videos.
- Experimental results validate the GPMK's superiority and show it outperforms state-of-the-art methods.
Conclusions:
- The GPMK offers a significant advancement in realistic human action recognition.
- Adaptive channel weighting provides a principled way to incorporate diverse information for improved accuracy.
- This approach enhances the robustness and effectiveness of kernel-based video classification systems.