Related Experiment Video
Updated: Apr 12, 2026

A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
Published on: January 18, 2020
Multimodal-guided prototype calibration and temporal coherence-aware hybrid matching for few-shot action recognition
Yiyuan An1, Yingmin Yi1, Yiwei Yuan1
1School of Automation and Information Engineering, Xi'an University of Technology, Xi'an, 710048, China.
None:
Few-shot action recognition methods achieve impressive performance by learning discriminative features and designing temporal alignment strategies. However, these methods suffer from three significant challenges in distinguishing similar classes: (a) spatiotemporal and motion features are complementary and essential, yet the latter is frequently overlooked; (b) excessive reliance on unimodal video data leaves multimodal information underexplored, and (c) the temporal distribution of actions is prone to intra-class temporal offsets and inter-class local similarity. To overcome these challenges, we propose a multimodal-guided prototype calibration and temporal coherence-aware hybrid matching (MGTH), which integrates four innovative components: a motion-enhanced temporal aggregation module (MTAM), a text-guided prototype calibration module (TPCM), a video-text adapter objective (VTA), and a temporal coherence-aware hybrid matching (TCH). The MTAM encodes complementary spatiotemporal and motion features without any 3D convolution or optical flow calculations. The TPCM fully utilizes video-text information to optimize video prototypes. Meanwhile, the VTA maximizes the similarity between video features and corresponding textual representations. The TCH strategy alleviates metric bias caused by intra-class temporal offsets and inter-class local similarity through a hybrid matching mechanism, and strengthens this mechanism's ability to distinguish similar actions by applying temporal coherence regularization to the input video. Additionally, we extend the proposed MGTH to more challenging tasks, including cross-domain few-shot action recognition and zero-shot action recognition. Experimental results on multiple benchmark datasets demonstrate that MGTH achieves state-of-the-art performance, confirming the superiority of our approach.