TS-D3D: A novel two-stream model for Action Recognition
Meng Yang, Yuanjun Guo, Feixiang Zhou, Zhile Yang · 2022 International Conference on Image Processing, Computer Vision and Machine Learning (ICICML) · 2022
Action recognition is a fundamental task of computer vision. The most widely used model in this task is Resnet. However, we want to study the capabilities of DenseNet in action recognition tasks. In this paper, we propose a new deep learning model named Two-Stream Densenet-3D (TS-D3D) for action recognition, taking the advantages of both 3D DenseNet and two-stream network. Specifically, we design an RGB pathway to extract spatial features and an RGB DIFF pathway to extract temporal features of continuous video frames, each of which is based on 3D DenseNet. In particular, we modify the convolution kernels of the RGB DIFF pathway for better modeling temporal correlations and reducing the computation of 3D convolution. Then, a Transition Layer is proposed to effectively fuse the features of the two pathways. Numerical results demonstrate that our TS-D3D performs well on widely adopted action recognition datasets.