Kinematics Features for 3D Action Recognition Using Two-Stream CNN
Jiangliu Wang, Yunhui Liu · 2018
Due to the great success of convolutional neural networks (CNN) on image classification problems, several attempts have been made to train deep neural networks for human action recognition problem. But since CNN is designed for static RGB images, it is not easy for it to learn temporal information from videos. To tackle this problem, temporal encoded kinematics features are proposed, which compute the linear velocity and orientation displacement based on human skeleton data. A two-stream CNN architecture is used, incorporating spatial and temporal networks. The spatial ConvNet is trained on still RGB images, while the temporal ConvNet is trained on the proposed encoded kinematics features. We evaluate our method on a popular and challenging 3D multiview human action benchmark, Northwestern-UCLA dataset. The experiment results show that our proposed method is fast to train and performs better when compared to traditional handcrafted features.