Multi-views Action Recognition on 3D ResNet-LSTM Framework

Jie Zhang, Fengshan Bai, Jianfeng Zhao, Song Zheng · 2021

In order to solve the problem of low accuracy rate of individual action recognition due to different camera views. In this paper, we propose a deep neural network framework based on 3D ResNet-LSTM for the recognition of individual actions of multiple views. The IXMAS dataset was chosen for the experiment, which was created from videos of 12 actions taken from 5 views. This dataset is preprocessed into a 5D tensor dataset and then used as the input to 3D ResNet. 3D ResNet can acquire both the appearance information of the image and the timing information between sequence frames. LSTM solves the problems of vanishing gradient and exploding gradient problems in the long time sequence training process. The two networks are connected in series through the reshape function, and the accuracy of image sequence recognition is significantly improved compared to other deep neural networks. Use the 3D ResNet-LSTM framework to train the 5D tensor dataset. Among the five views, the front view and the right oblique upper view had better recognition, with accuracy rates of 93.7% and 92.3%, respectively, and the overhead view had the lowest accuracy rate. By analyzing the accuracy of the five views, we fuse the data of the image sequences from the front-front view and the right oblique top view and produces a 5D tensor dataset, which finally yields an accuracy of 94.9%, which is higher than the accuracy of the individual views.

Read the paper · More papers on PaperTik