Modality-convolutions: Multi-modal gesture recognition based on convolutional neural network

Da Huo, Yufeng Chen, Fengxia Li, Zhengchao Lei · 2017

We proposed a novel method of feature extraction for multi-modal images called modality-convolution. It extracts both the intra- and inter-modality information. Whats more, it completes the data fusion at pixel-level so that the complementarity of information contained in multi-modal data is fully utilized. Based on the modality-convolution, we describe a modality-CNN for multi-modal gesture recognition. For extracting the features in RGB-D images, the modality-CNN is adopted in the gesture recognition framework. The framework use DBN to present the skeleton data. Then, the probability obtained by the two networks are fused and put into the HMM to carry out dynamic gestures classification. We use the Jaccar Index to calculate the accuracy of gesture recognition. A comparative experiment on ChaLearn LAP 2014 gesture datasets shows that the modality-convolution is able to extract the inter- and intra-modality information effectively, which is helpful to improve the accuracy.

Read the paper · More papers on PaperTik