Related Experiment Video
Updated: Jul 10, 2025

05:41
A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
Published on: February 6, 2020
9.4K
ActionCLIP: Adapting Language-Image Pretrained Models for Video Action Recognition
IEEE Transactions on Neural Networks and Learning Systems
|November 21, 2023
Summary
This study introduces a new multimodal learning framework for video action recognition, treating it as a video-text matching problem. This approach enables zero-shot recognition and improves performance on unseen concepts.
Area of Science:
- Computer Vision
- Machine Learning
- Natural Language Processing
Background:
- Traditional video action recognition models are limited by fixed category predictions.
- These models lack transferability to new datasets with unseen concepts.
Purpose of the Study:
- To propose a novel approach for video action recognition using multimodal learning.
- To enable zero-shot action recognition by leveraging semantic information from label texts.
- To develop a new "pre-train, adapt and fine-tune" paradigm for enhanced action recognition.
Main Methods:
- Modeling action recognition as a video-text matching problem within a multimodal framework.
- Utilizing semantic information from label texts for stronger video representation.
- Implementing a "pre-train, adapt and fine-tune" paradigm using large-scale web data.
- Introducing ActionCLIP as an instantiation of the proposed paradigm.
Main Results:
- Achieved superior and flexible zero-shot/few-shot transfer ability.
- Reached top performance on general action recognition tasks.
- Obtained 83.8% top-1 accuracy on Kinetics-400 using a ViT-B/16 backbone.
Conclusions:
- The proposed multimodal learning framework significantly enhances video action recognition.
- The "pre-train, adapt and fine-tune" paradigm offers a powerful approach for action recognition.
- ActionCLIP demonstrates strong performance and excellent transferability for zero-shot and few-shot learning.

