Multilevel Fusion with Dual Stream 3DCNN-LSTM for Advancing Dynamic Hand Gesture Recognition
Bristy Chanda, Hussain Nyeem · 2023
This paper introduces a novel method for identifying dynamic hand gestures in multimedia applications, tackling the task of analysing time-varying characteristics within video streams. Our proposed method utilizes a multi-level fusion network that combines RGB and depth modalities using Long Short-Term Memory (LSTM) and a 3D Convolutional Neural Network (3DCNN). Unlike existing methods, we incorporate a U-Net-based semantic segmentation technique to extract depth features via a gesture-specific mask. The 3DCNN and LSTM models are employed to extract spectral and spatial features from both RGB and depth images. We evaluate our approach using the 20BN-Jester dataset (ten classes) and the Sebastien Marcel Dynamic Hand Posture dataset (four classes). Our results indicate that our late fusion model, which combines depth and RGB data, achieves an average validation accuracy of 97.8% on the 20BN-Jester dataset and 98.5% on the Sebastien Marcel Dynamic Hand Posture dataset, demonstrating its effectiveness in dynamic hand gesture recognition.