Action Recognition Based on Spatial Temporal Graph Convolutional Networks

Wanqiang Zheng, Punan Jing, Qingyang Xu · Proceedings of the 3rd International Conference on Computer Science and Application Engineering · 2019

Compared with the achievements of convolutional neural networks in image classification, human action recognition for video is not ideal in terms of accuracy and practicability. A major method in action recognition is based on the human skeleton, which is an important information for characterizing human motion in video. In this paper, the human skeleton in video is extracted by OpenPose, and the spatial and temporal graph of skeleton is constructed. The spatial and temporal graph convolution network (ST-GCN) is used to extract the spatial and temporal features of the human skeleton on consecutive video frames, and the features is used for video classification. In order to verify the action recognition performance based on the ST-GCN, a 50.53% top-1 and 81.58% top-5 accuracy is obtained on the UCF-101 dataset. A specific UCF-31 dataset is constructed manually and a 68.73% top-1 and 94.43% top-5 accuracy is obtained, verifying that the identification accuracy of ST-GCN model would also be improved when the accuracy of skeleton acquisition was improved.

Read the paper · More papers on PaperTik