Action Recognition Networks Based on Spatio-Temporal Motion Modules
Yimin Zhang, Huifang Qian, Jialun Zhang, Zhenyu Shi · 2024
The spatiotemporal feature is of paramount importance in the area of video action recognition. Previously, typical methods employed 3D CNN to process both spatial and temporal features, but this approach is computationally expensive. Alternative methodologies utilise (1+2)D CNNs to efficiently learn spatiotemporal features, However, they fail to acknowledge the significance of motion representation. In order to surmount the problems of previous methods, In this paper, a new STM module is proposed to alleviate the above problems, which is able to capture spatiotemporal features more realistically and learn motion features efficiently. STM module includes Motion Excitation (ME), Multi-view Spatio-Temporal Excitation (MSTE). The ME channel computes feature-level time differences from spatiotemporal features and then uses the time differences to excite the motion sensitivity of the feature channels; The MSTE channel performs feature extraction for spatial and temporal dimensions along the three orthogonal views in a consistent manner, obtaining useful information in the video action from the different views. In this paper, the STM module was inserted into 2D ResNet-50 and extensive experiments were carried out on the datasets (Something-Something v2, UCF101 and HMDB51) separately, respectively, and the results show that it outperforms the previous CNN-based methods in terms of the accuracy of "Val Top-1%" and "Val Top-5%". The STM module proposed in this paper is an efficient method for extracting spatio-temporal and motion features, which significantly enhances the performance of the action recognition baseline through the use of a lightweight adaptive embedding.