Improving Dynamic Hand Gesture Recognition on Multi-views with Multi-modalities
Huong-Giang Doan, Van-Toi Nguyen · International Journal of Machine Learning and Computing · 2019
Hand gesture recognition topic has been researched for many recent decades because it could be used in many fields as sign language, virtual game, human-robot interaction, entertainment and so on.However, this problem has been faced to many challenges such as combination of multi information in a temporal flow in order to understand the meaning of human hand gesture.In recent times, thanks to the advances in hardware technologies such as readily available 3D cameras, Kinect sensors, and etc.The impressive performance of cutting-edge techniques in computer vision, which is known as: manifold learning, deep learning techniques and/or the presentation of various multimodal fusion strategies.There have been many improvements in exploiting of features from multimodal data to effectively solve human hand gesture recognition tasks.Therefore, this paper focuses on solving the problem of dynamic hand gesture recognition in our daily life.We consider methods for extracting features of different data sources (RGB images and depth images) based on both manifold learning and deep learning technique.For RGB information, a manifold technique is performed to extract spatial feature that is then composed with temporal feature extracted by KLT technique.Among many deep learning architectures proposed in the literature that achieved good results in detecting human activities, I studied and proposed a simple convolutional neural network to extract feature of depth motion map.This technique extracts hand features from depth information which combines spatial and temporal aspects.Besides that, fusion algorithms are deployed to unite with those extracted features and enhance the accuracy of a final dynamic hand gesture results.Evaluation results confirm that the best accuracy rate achieves at 84.7% that is significantly higher than results from previous works (at 78.4%).The proposed method suggests a feasible solution addressing technical issues in using multimodality and multi-viewpoint of hand gestures.