Related Experiment Video
Updated: Jun 1, 2025

04:43
Visualizing Visual Adaptation
Published on: April 24, 2017
8.9K
A discriminative multi-modal adaptation neural network model for video action recognition
Summary
This study introduces a novel two-stream heterogeneous network for multi-modal action recognition, effectively integrating RGB and skeleton data. The proposed discriminative multi-modal adaptation neural network model (DMANNM) significantly improves recognition accuracy by leveraging statistical machine learning and convolutional neural networks.
Area of Science:
- Computer Vision
- Machine Learning
- Artificial Intelligence
Background:
- Video-based understanding is crucial for applications like e-healthcare and affective computing.
- Multi-modal action recognition faces challenges in effectively utilizing cross-modality information.
- Existing score-level fusion methods for multi-modal action recognition yield sub-optimal performance due to inadequate cross-modality semantic consideration.
Purpose of the Study:
- To address the limitations of current multi-modal action recognition techniques.
- To propose a novel approach for effectively extracting and jointly processing complementary features from different modalities.
- To enhance the performance of video-based action recognition systems.
Main Methods:
- A two-stream heterogeneous network is developed to extract features from RGB and skeleton modalities.
- A discriminative multi-modal adaptation neural network model (DMANNM) is proposed, integrating statistical machine learning (SML) with convolutional neural network (CNN) architecture.
- An effective nonlinear classification algorithm is employed to achieve high recognition accuracy.
Main Results:
- The proposed DMANNM model demonstrates superior performance compared to state-of-the-art methods.
- Experiments conducted on four diverse datasets (NTU RGB+D, NTU RGB+D 120, N-UCLA, SYSU) validate the model's effectiveness and generalizability.
- The integrated SML and CNN architecture provides an adaptive platform for handling datasets of varying scales.
Conclusions:
- The developed two-stream heterogeneous network and DMANNM model offer a significant advancement in multi-modal action recognition.
- The proposed method effectively overcomes the limitations of score-level fusion by considering cross-modality semantics.
- The model's adaptability and high accuracy make it a promising solution for real-world video-based action recognition tasks.

