Related Experiment Video
Updated: Jan 12, 2026

12:39
A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
Published on: January 18, 2020
8.1K
${\text{CA}^{2}\text{ST}}$: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition
IEEE Transactions on Pattern Analysis and Machine Intelligence
|November 4, 2025
Summary
Cross-Attention in Audio, Space, and Time (CA²ST) enhances video recognition by integrating spatial, temporal, and audio information. This transformer-based method achieves balanced and holistic video understanding through synergistic expert information exchange.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Action recognition in videos necessitates understanding both spatial and temporal dynamics.
- Existing models often struggle to achieve a balanced spatio-temporal comprehension of video content.
Purpose of the Study:
- To introduce a novel transformer-based method for holistic video recognition.
- To develop a model that achieves balanced spatio-temporal understanding and integrates audio-visual information.
Main Methods:
- Proposed Cross-Attention in Space and Time (CAST), a two-stream RGB-input architecture.
- Introduced Bottleneck Cross-Attention (B-CA) for synergistic information exchange between spatial and temporal experts.
- Extended CAST to CAVA by integrating an audio expert for audio-visual recognition.
- Developed CA²ST, combining CAST and CAVA for multi-expert cross-attention.
Main Results:
- CAST demonstrated balanced performance across diverse video recognition benchmarks (EPIC-KITCHENS-100, Something-Something-V2, Kinetics-400, ActivityNet, HD-EPIC).
- CAVA achieved favorable results on audio-visual action recognition datasets (UCF-101, VGG-Sound, KineticsSound, EPIC-SOUNDS, HD-EPIC-SOUNDS).
- The B-CA module effectively facilitated information exchange among multiple experts.
Conclusions:
- CA²ST provides a balanced and holistic approach to video understanding by effectively integrating spatial, temporal, and audio information.
- The proposed cross-attention mechanism enables synergistic predictions from diverse expert modules.

