A Dynamic Gesture Recognition Method Based on MobileNetV4
Hailun Du, Diansheng Chen, Xiaochuan Zhang, Yazhe Luo, Yifei Li, Haosong Ran · 2024
Dynamic gesture recognition, as a convenient human-computer interaction method, is widely applied in various fields. Although current deep learning-based dynamic gesture recognition models have achieved high accuracy, the pursuit of even higher accuracy has led to increasingly deeper networks and higher model complexity. This, in turn, demands significant hard-ware resources, making it challenging to deploy on embedded devices and mobile platforms. Therefore, we build a lightweight dynamic gesture recognition model with good recognition effect based on MobileNetV4. We propose a method to train continuous image data for dynamic gesture recognition using the 2D version of MobileNetV4. We validate the limitation of 2D CNN in extracting temporal information features when handling sequential data. Then, we introduce 3D CNN, which has the capability to extract temporal information. Based on the MobileNetV4 architecture, we construct two 3D versions with different complexities, referred to as 3D-MobileNetV4. We evaluate the performance of 3D-MobileNetV4 against several mainstream classification networks using four evaluation metrics on the public Jester dataset. The more complex 3D-MobileNetV4-Medium achieved the highest recognition accuracy among all comparison models, reaching 90.23%, with relatively low computational and parameter costs. The simpler 3D-MobileNetV4-Small achieved the highest accuracy among all compared lightweight models, reaching 87.93%, with lower computational requirements than most lightweight models. Results demonstrate that the proposed 3D-MobileNetV4 can be adapted to different practical applications, allowing the selection between two different complexity levels of 3D-MobileNet to accomplish dynamic gesture classification tasks based on specific needs.