Improving Static Hand Gesture Recognition with Semantic Segmentation and Fine-Tuned Convolutional Neural Network
Bristy Chanda · 2024
The study aims to develop an effective static gesture recognition system for RGB datasets, enabling hand-based interaction between users and virtual environments. Using the U-Net architecture, ROI is first extracted from the RGB input data. Multilevel thresholding, morphological filters, and logical AND operation are then used during processing to enhance depth-segmented images. The small amount of levelled image samples in the static hand gesture database makes it difficult to train popular CNN networks like as VGG16, VGG19, ResNet50, and Inceptionv3 from scratch. Consequently, an end-to-end fine-tuning strategy of a pre-trained CNN model with score-Level fusion technique is proposed here to recognise hand gestures in a dataset with a low number of gesture photos, inspired by CNN performance. The proposed model is evaluated on the National University of Singapore (NUS) hand posture dataset II (subset A), which contains 2000 images in 10 classes, compared with a few pre-trained CNN architectures like VGG16, VGG19, and ResNet-50, Inception V3. The results of the experiments illustrate that the proposed approach, which combines the U-Net architecture for semantic segmentation and score level fusion between fine-tuned ResNet50 and VGG16 for feature extraction and classification, performs better than alternative approaches.