An Effective Video Transformer With Synchronized Spatiotemporal and Spatial Self-Attention for Action Recognition

Summary

We introduce three techniques to enhance video understanding using video Transformers, improving efficiency and performance. Our novel methods, including synchronized spatiotemporal and spatial attention, achieve state-of-the-art results on benchmark datasets.