Three-stream Very Deep Neural Network for Video Action Recognition

Nasim Khani, Mehdi Rezaeian · 2019

The purpose of this study is to determine whether fine-tuning very deep three-dimensional Convolutional Neural Network (3D CNN) that already pre-trained on an adequately large video dataset will give sufficient motion information for action recognition or still need to have supplementary information. We introduce a three-stream CNN that is based on two-dimensional (2D) and 3D kernels while leveraging successful pre-trained networks on ImageNet and Kinetics datasets. In order to analyze these streams, we fine-tune each on the HMDB-51 challenging dataset and show that supplementary motion information (optical flow and the proposed sparse trajectory image) are critical to action recognition despite using 3D CNN. Experimental outcomes determine that our network reaches 80.92% accuracy on the HMDB-51 dataset and its performance is comparable with the performance of state-of-the-art networks on this dataset.

Read the paper · More papers on PaperTik