Video swin-CLSTM transformer: Enhancing human action recognition with optical flow and long-term dependencies

Jun Qin1,2, Shenwei Chen1,3, Zheng Ye1

  • 1College of Computer Science, South-Central Minzu University, Wuhan, Hubei, China.

Plos One
|July 7, 2025
PubMed
Summary

This study introduces the Video Swin-CLSTM Transformer for improved Human Action Recognition (HAR). The model enhances accuracy and efficiency by integrating optical flow and ConvLSTM, outperforming existing methods on the UCF-101 dataset.