Related Experiment Videos
Interactive image-to-video transfer learning
Cong Wu1, Tianyang Xu2, Zhenhua Feng2
1School of Artificial Intelligence and Computer Science, Jiangnan University, 214122, China; Postdoctoral Research Station in Design, Jiangnan University, 214122, China.
Abstract:
Transfer learning from image to video has become a widely adopted strategy in action recognition. Existing mainstream approaches typically fine-tune the entire network after initialising with pre-trained parameters, which inevitably leads to substantial training overhead. Recent efficiency-oriented methods alleviate this by freezing the backbone during optimisation; however, they tend to overly rely on powerful prior knowledge encoded in image models and often neglect how to effectively perform video reasoning with static backbones. In this paper, we propose an efficient image-to-video transfer learning framework, termed SDST (Static-Dynamic & Spatial-Temporal interaction framework), which explicitly enhances the interaction between static and dynamic cues as well as spatial and temporal domains. Specifically, we introduce the Motion Booster Module that extracts motion descriptors from intermediate features and integrates them via the progressive cross-attention mechanism to fuse static and dynamic representations. Furthermore, we propose a lightweight temporal modelling module, named Channel-Aware Multi-Scale Temporal Modelling, characterised by channel awareness, various temporal scales modelling, as well as feature interactions. By enriching the frozen image backbone through these components, our approach effectively bridges the gap between static visual representations and video-based action recognition tasks. Extensive experiments on Something-Something V1&V2, Diving-48, and Kinetics-400 demonstrate that our method not only surpasses State-of-the-Art efficient transfer learning techniques but also outperforms several fully fine-tuned approaches, all without additional bells and whistles. Moreover, we validate the transferability of the proposed SDST framework to vision-language pre-trained models like CLIP, achieving further gains in performance. These results highlight the potential of our method as a general and scalable solution for efficient video understanding in the era of large vision-language models.
Related Concept Videos
Observational Learning
Introduction to Learning
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
Deconvolution
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Vision
Transformation
Improving Translational Accuracy