Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking.

IEEE transactions on neural networks and learning systems·2025
Same author

Heterogeneity of late endosome/lysosomes shown by multiplexed DNA-PAINT imaging.

The Journal of cell biology·2024
Same author

Camera-Based 3D Semantic Scene Completion With Sparse Guidance Network.

IEEE transactions on image processing : a publication of the IEEE Signal Processing Society·2024
Same author

Multiplexed DNA-PAINT Imaging of the Heterogeneity of Late Endosome/Lysosome Protein Composition.

bioRxiv : the preprint server for biology·2024
Same author

Interplay between stochastic enzyme activity and microtubule stability drives detyrosination enrichment on microtubule subsets.

Current biology : CB·2023
Same author

Delving Deeper Into Mask Utilization in Video Object Segmentation.

IEEE transactions on image processing : a publication of the IEEE Signal Processing Society·2022

Related Experiment Video

Updated: Jul 10, 2025

A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
05:41

A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis

Published on: February 6, 2020

9.4K

ActionCLIP: Adapting Language-Image Pretrained Models for Video Action Recognition.

Mengmeng Wang, Jiazheng Xing, Jianbiao Mei

    IEEE Transactions on Neural Networks and Learning Systems
    |November 21, 2023
    PubMed
    Summary

    This study introduces a new multimodal learning framework for video action recognition, treating it as a video-text matching problem. This approach enables zero-shot recognition and improves performance on unseen concepts.

    More Related Videos

    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
    08:25

    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

    Published on: May 7, 2019

    9.0K
    Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
    03:14

    Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

    Published on: December 6, 2024

    596

    Related Experiment Videos

    Last Updated: Jul 10, 2025

    A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
    05:41

    A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis

    Published on: February 6, 2020

    9.4K
    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
    08:25

    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

    Published on: May 7, 2019

    9.0K
    Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
    03:14

    Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

    Published on: December 6, 2024

    596

    Area of Science:

    • Computer Vision
    • Machine Learning
    • Natural Language Processing

    Background:

    • Traditional video action recognition models are limited by fixed category predictions.
    • These models lack transferability to new datasets with unseen concepts.

    Purpose of the Study:

    • To propose a novel approach for video action recognition using multimodal learning.
    • To enable zero-shot action recognition by leveraging semantic information from label texts.
    • To develop a new "pre-train, adapt and fine-tune" paradigm for enhanced action recognition.

    Main Methods:

    • Modeling action recognition as a video-text matching problem within a multimodal framework.
    • Utilizing semantic information from label texts for stronger video representation.
    • Implementing a "pre-train, adapt and fine-tune" paradigm using large-scale web data.
    • Introducing ActionCLIP as an instantiation of the proposed paradigm.

    Main Results:

    • Achieved superior and flexible zero-shot/few-shot transfer ability.
    • Reached top performance on general action recognition tasks.
    • Obtained 83.8% top-1 accuracy on Kinetics-400 using a ViT-B/16 backbone.

    Conclusions:

    • The proposed multimodal learning framework significantly enhances video action recognition.
    • The "pre-train, adapt and fine-tune" paradigm offers a powerful approach for action recognition.
    • ActionCLIP demonstrates strong performance and excellent transferability for zero-shot and few-shot learning.