Convolutional LSTM Networks and RGB-D Video for Human Motion Recognition

Weisong Che, Shuhua Peng · 2018 IEEE 4th Information Technology and Mechatronics Engineering Conference (ITOEC) · 2018

In this paper, we present a novel deep neural network architecture for addressing the problem of human motion recognition. The proposed architecture enables direct mapping of RGB-D input video to output classification through an end-to-end approach. The RBG images and the depth images pass through a two-stream deep convolutional long short-term memory networks, thereby outputting a plurality of feature maps including complete spatiotemporal features of the entire sequence. Then, a couple of deep convolutional neural networks further extracts the high-level features of these feature maps and classification of action through a classifier. This network structure allows the spatiotemporal structure of the image sequence to be preserved entirely from start to finish and improves the robustness of the network through the separate processing of RGB information and depth information. The experimental results show that the recognition accuracy of this structure in the SBU Kinect Interaction Dataset and MSRDailyActivity3D Dataset is higher than other existing structures.

Read the paper · More papers on PaperTik