Related Experiment Video
Updated: Aug 14, 2026

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
SASA-CLIP: Structure-Aware Alignment with a Gaussian Prior for Fine-Grained Video Action Recognition
Xiaowei Han1,2, Wenbao Si1,2, Honghui Zhang1,2
1School of Computer and Information Engineering, Harbin University of Commerce, Harbin 150028, China.
We introduce Structure-Aware Semantic-Adaptive CLIP (SASA-CLIP) for fine-grained video action recognition. SASA-CLIP improves accuracy by aligning multi-granular features with a temporal structure prior, outperforming existing methods.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Fine-grained video action recognition is challenging due to subtle inter-class variations and complex temporal dynamics.
- Existing Contrastive Language-Image Pre-training (CLIP) methods often use early global pooling, losing crucial temporal details for subtle action differentiation.
- This leads to a granularity mismatch in cross-modal alignment, hindering performance on detailed action recognition tasks.
Purpose of the Study:
- To propose a novel framework, Structure-Aware Semantic-Adaptive CLIP (SASA-CLIP), for effective multi-granular cross-modal alignment in video action recognition.
- To address the limitations of early global pooling in conventional CLIP-based approaches by preserving fine-grained temporal information.
- To enhance the accuracy and robustness of fine-grained video action recognition through a structure-aware alignment process.
Main Methods:
- SASA-CLIP employs a dual-branch design: a coarse-grained branch for global context and a fine-grained branch for frame-level descriptor matching before aggregation.
- A Gaussian prior is introduced as a temporal structural constraint, encoding local temporal continuity into the attention matrix.
- This guides an ordered alignment of key action segments along the temporal axis, ensuring temporal coherence.
Main Results:
- SASA-CLIP achieved a Top-1 accuracy of 81.37% on Kinetics-400 (ViT-B/32), surpassing the X-CLIP baseline by 0.97%.
- On HMDB-51 and UCF-101 datasets (ViT-B/16), SASA-CLIP reached 74.0% and 96.81% accuracy, respectively, showing improvements of 3.25% and 2.61%.
- The framework demonstrated improved performance in the zero-shot setting on HMDB-51 and UCF-101, indicating strong generalization capabilities.
Conclusions:
- Combining multi-granular representations with a temporal structural prior significantly benefits fine-grained video action recognition.
- SASA-CLIP offers a practical solution for real-world visual sensing applications, including intelligent surveillance and wearable activity monitoring.
- The proposed method effectively bridges the granularity gap in cross-modal alignment for detailed action understanding.
Related Concept Videos
Absolute Motion Analysis- General Plane Motion
As the drone's propellers rotate, an upward force is generated that counteracts the force of gravity, enabling the drone to lift off from the ground. This initial movement of the drone is along a straight path, representing a form of translational motion. In this phase, every point on the drone...
Relative Motion Analysis using Rotating Axes
However, to express the relative position of point B relative to point A, an additional frame of reference, denoted as x'y', is necessary. This additional frame not only translates but also rotates relative to the fixed frame, making it instrumental in...
Fixed Action Patterns
Relative Motion Analysis - Acceleration
Relative Motion Analysis using Rotating Axes-Problem Solving
Here, in order to determine the magnitude of velocity and acceleration for point...
Curvilinear Motion: Rectangular Components
As the car advances, its position evolves over time. Quantifying the car's velocity involves computing the time...