Related Experiment Video
Updated: Sep 22, 2025

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
9.1K
MMNet: A Model-Based Multimodal Network for Human Action Recognition in RGB-D Videos
Summary
This study introduces a novel multimodal network (MMNet) for human action recognition (HAR) using RGB-D videos. MMNet effectively fuses skeleton and RGB data, outperforming existing methods on multiple datasets.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Human Action Recognition (HAR) in RGB-D videos is a growing field.
- Unimodal approaches (skeleton-based, RGB-based) have advanced significantly.
- Multimodal methods, especially model-level fusion, remain underexplored.
Purpose of the Study:
- To propose a model-based multimodal network (MMNet) for fusing skeleton and RGB data.
- To enhance ensemble recognition accuracy by leveraging complementary information from different modalities.
- To improve the discriminative power of features for HAR.
Main Methods:
- Developed a model-based multimodal network (MMNet).
- Employed a spatiotemporal graph convolution network for skeleton data to learn attention weights.
- Transferred learned attention weights from skeleton to RGB modality network.
- Utilized model-level fusion to combine skeleton and RGB information.
Main Results:
- MMNet outperformed state-of-the-art approaches on six evaluation protocols across five benchmark datasets (NTU RGB+D 60/120, PKU-MMD, Northwestern-UCLA, Toyota Smarthome).
- The method demonstrated consistent performance on the Kinetics 400 RGB video dataset, indicating robustness.
- Achieved effective capture of complementary features between RGB and skeleton modalities.
Conclusions:
- The proposed MMNet effectively fuses skeleton and RGB modalities for improved HAR.
- MMNet provides more discriminative features by capturing complementary information.
- The model-level fusion approach offers a promising direction for multimodal HAR research.

