ActionCLIP: Adapting Language-Image Pretrained Models for Video Action Recognition

Summary

This study introduces a new multimodal learning framework for video action recognition, treating it as a video-text matching problem. This approach enables zero-shot recognition and improves performance on unseen concepts.

Related Concept Videos