Gesture Recognition Based on Improved AlexNet
Renbo Liu, Aimin Xiong, Jinghao Lai, Hongbin Zhang, Suqin Wu · 2022
Gesture recognition is one of the research hotspots in the field of computer science and technology today, aiming to establish a natural and harmonious human-computer interaction environment and a colorful interaction mode. The application range of gesture recognition is very wide: gesture recognition can not only help deaf people to solve interaction problems, but also be applied in various fields such as daily life, education, business office, etc. Meanwhile, gesture recognition technology is also needed in the more popular metaverse recently. Handcrafted methods for vision-based gesture recognition usually involve multiple stages of professional processing, such as traditional handcrafted-based feature extraction methods, which are generally used to specialize in specific tasks with insufficient generalization capability. For this reason, this paper proposes an improved AlexNet-based gesture recognition algorithm by stacking four 3×3 convolutional kernels instead of the 11×11 convolutional kernels in AlexNet. In addition, the number of training data is increased and its diversity is enriched by preprocessing such as random cropping and random level flipping to further improve the generalization ability of the algorithm. The improved algorithm proposed in this paper further enhances the feature extraction capability of the network, and the experiments are conducted on multiple datasets, i.e., one American Sign Language dataset and two NUS gesture datasets. The average accuracy of the improved algorithm is 95.3%, which is better than the original AlexNet network as well as other deep learning algorithms.