Skeleton-Based View Invariant Deep Features for Human Activity Recognition

Chhavi Dhiman, Manan Saxena, Dinesh Kumar Vishwakarma · 2019

Recently skeleton-based human activity recognition is receiving significant attention due to its efficient localization of human pose in 3D space. However, the state-of-the-arts are utilizing majorly the temporal information of the skeleton joint coordinate values and few transformed the skeleton coordinate values into images to process spatial patterns. These images needs to be up-sampled to the size of the CNN models that introduces noise. This paper introduces a novel view-invariant human action recognition framework by using skeleton joints coordinates which exploit spatial shape structures of the skeletons and RGB Dynamic Images (DIs). Skeleton coordinates based features help to localize human pose and its variations in orientation during the action performed and reduces the noise element when fed to CNN models. Whereas, DIs [1] encrypt the temporal dynamics of action by using the concept of Average Rank pooling. This hybrid representation of human features is projected into higher dimensional space by applying the concept of transfer learning on InceptionV3 [1] architecture, which is fine-tuned for two multi-view human action datasets. In addition, the final prediction of the action class is made by applying late fusion on skeleton-based features and DIs features. The performance of the proposed scheme is tested for two multi-view NUCLA and NTU RGB+D Dataset.

Read the paper · More papers on PaperTik