Shift Swin Transformer Multimodal Networks for Action Recognition in Videos
Xiongjiang Xiao, Ziliang Ren, Wenhong Wei, Huan Li, Hua Tan · 2022
Action recognition aims to understand human behavior. Currently, action recognition is mainly studied using deep neural networks. The traditional 2D CNN (Convolutional Neural Network) has low calculation cost, but it is difficult to capture action sequence information; the 3D CNN-based methods can extract excellent spatiotemporal information, but it has many parameters and a large amount of computation. To improve accuracy and robustness, we propose a Shift Swin Transformer Multimodal Networks method. The shift design is adapted from the Temporal Shift Module (TSM), which facilitates information exchange between adjacent frames by shifting partial channels along the temporal dimension. Furthermore, based on the superiority of Swin Transformer network and TSN (Temporal Segment networks), a feature learning method is proposed to improve the performance. Extensive experiments demonstrate that our method markedly improves the effect in video action recognition, attaining 79.8% on HMDB-51, 97.4% on UCF-101 and 76.5% on Kinetics-400. Besides, we get a cross-subject accuracy of 88.6% and a Cross-Setup accuracy of 89.7%on NTU RGB+D 120.