Towards 3D Human Action Recognition Using a Distilled CNN Model
Jingjing Ren, Napoleon H. Reyes, Andre L. C. Barczak, Chris J. Scogings, Mingzhe Liu · 2018
Motivated by the remarkable performance achieved using deep learning strategies in solving action recognition tasks, an effective, yet simple method is proposed for encoding the spatiotemporal information of skeleton sequences into color texture images, referred to as Skeletal Optical Flows (SOFs). SOFs collectively represent the kinetic energy, predefined angles and pair-wise displacements between joints over consecutive frames of skeleton data, as color variations to capture meaningful temporal information and make them highly interpretable. A novel Convolutional Neural Network with Correctness-Vigilant Regularizer (CVR-CNN) is then employed to exploit the discriminative features of SOFs for human action recognition. Empirical results show that the efficiency of the proposed method is superior in terms of the generalizability of the generated model, the training convergence speed, and the resulting classification accuracy on commonly used action recognition datasets, such as MHAD, HDM05 and NTU RGB+D.