From Coarse to Fine: Hierarchical Multi-scale Temporal Information Modeling via Sub-group Convolution for Video Action Recognition
Fengwen Cheng, Huicheng Zheng, Zehua Liu · 2021
In the video action recognition task, it is essential to model the temporal information. Since different actions have different durations, capturing multi-scale temporal features is very crucial. In this paper, we propose a multi-scale modeling (MSM) module to exploit temporal information for action recognition, which is composed of a multi-scale temporal convolution (MTC) block and a multi-scale hierarchical convolution (MHC) block. MTC uses convolutions of multiple temporal depths to capture features of different temporal scales to enhance the connection between frames. MHC employs group convolution to obtain more fine-grained multi-scale features in the channel dimension. In MHC, the convolutions are formulated as a hierarchical structure to expand the receptive fields as well as help implement a series of sub-group convolutions, which can help to realize the modeling of distinctive and long-term information. The two components of MSM are complementary in temporal modeling. Finally, we evaluated our method on several action recognition benchmarks including Kinetics, UCF10l, and HMDB51, and obtained competitive results, which verified the effectiveness of the proposed method in temporal modeling,