A Lightweight-Grouped Model for Complex Action Recognition

Bingkun Gao, Yunze Bi, Hongbo Bi, Le Dong · Pattern Recognition and Image Analysis · 2021

Abstract With the rise of video analysis, temporal reasoning is playing an extremely important part, also the recognition of complex action is turning into a new challenge. 3D convolution neural networks (CNN) can better explore the spatial-temporal features, but it also increases the computational cost. Although the 2D CNN has small parameters, it lacks of key termporal reasoning, as a result, the effect of 2D CNN is not fully achieved. In this paper, we design a new lightweight-grouped (LWG) model. On the basis of the decomposition of feature channels, we decompose the 3D convolution kernel into two spatial-temporal convolution kernels in horizontal and vertical directions, and, then, we use the depthwise separable convolution to replace the common convolution. This method is more parameter-efficient and solves the problem of computational cost which explore how the temporal channel capacity affects spatial-temporal modeling. We validate our model on the action recognition dataset and demonstrate its effectiveness, the results show that this method has fewer parameters than convolutional 2D network (C2D) and better performance than convolutional 3D network (C3D). It is also very competitive compared with the current mainstream models.

Read the paper · More papers on PaperTik