Action Recognition with 3D ConvNet-GRU Architecture
Guangle Yao, Xianyuan Liu, Tao Leí · 2018
Video action recognition is widely applied in video indexing, intelligent surveil-lance, multimedia understanding, and other fields. Recently, it was greatly improved by incorporating the learning of deep information using convolutional neural network (ConvNet). In this paper, we proposed a 3D ConvNet-GRU architecture to learn deep information for action recognition. Specifically, we use 3D ConvNet to learn spatiotemporal information from short RGB clips and optical flow clips, and impose gated recurrent unit (GRU) on the spatiotemporal in-formation to model the temporal evolution for action recognition. The experimental results show that our 3D ConvNet-GRU method is effective to model temporal evolution for action and achieves recognition performance comparable to that of state-of-the-art methods.