Related Experiment Videos

ViT-ConvGAN: a hybrid model for spatiotemporal action recognition using video transformer and 3D CNN

Chenlei Miao1,2, Jianjun Lin2, Lin Jia3

  • 1School of Physical Education, Guangzhou University, Guangzhou, 510006, Guangdong, China.

Scientific Reports
|June 12, 2026
PubMed
Summary

This study introduces ViT-ConvGAN, a novel model for video action recognition. It effectively balances global and local motion understanding, achieving high accuracy on complex actions.

Related Concept Videos