Related Experiment Videos
Exploring Multimodal Adapters with Multi-Task Decoders for Video Action Recognition
Abstract:
Large-scale vision-language pretrained models like CLIP, paired with parameter-efficient fine-tuning techniques, have emerged as promising solutions for image-to-video transfer in video action recognition. However, existing methods often prioritize strong supervised performance at the expense of transferability and generalization. To address this challenge, we propose a novel multimodal, multi-task CLIP adaptation framework, M$^{2}$-CLIP, which effectively balances high supervised performance with robust generalization. Our approach begins by enhancing individual modality architectures with multimodal adapters tailored for both the visual and textual branches. Specifically, we propose TED-Adapter, which improves temporal modeling by capturing both global temporal enhancements and local temporal differences within the visual encoder. We also integrate adapters into the text encoder to strengthen the semantic alignment between visual and textual modalities. To further boost performance, we incorporate a multi-task decoder that leverages diverse supervisory signals, including contrastive learning, cross-modal classification, masked language modeling, and visual classification tasks. This multi-task decoder ensures robust supervised performance while preserving multimodal consistency. Building on the success of M$^{2}$-CLIP, we introduce M$^{2}$-CLIP++, which enhances I2V feature modeling through advanced visual feature integration. Conventional fusion strategies typically adopt a sequential paradigm or rely on elementary operations (e.g., addition), which fail to capture the complex, non-linear correlations between spatial appearance and temporal dynamics. To address this limitation, M$^{2}$-CLIP++ is explicitly designed to enable higher-order temporal and spatial feature fusion, allowing the model to better interpret how spatial elements evolve over time. Specifically, it updates the visual adapter to TED-Adapter++, which employs a dual learnable bilinear fusion strategy to achieve this sophisticated interaction modeling. Combined with multi-head processing and low-rank decomposition, our approach optimizes computational efficiency while maintaining effective representation capacity. Additionally, M$^{2}$-CLIP++ incorporates multi-scale feature aggregation to effectively merge features across different semantic levels. Extensive experiments demonstrate that our models achieve state-of-the-art supervised learning performance while maintaining exceptional generalization capabilities. Code will be made available at https://github.com/sallymmx/m2clip.