A Fusion of CNN, MLP, and MediaPipe for Advanced Hand Gesture Recognition

Kulsuma Akter Priyanka, R. R. Rithik, Chamarthi Jaswanth, Cherukupalli Rajesh · 2024

Hand gesture recognition has become evident as a critical component in human-centered computing systems. It provides users a natural intuitive way for users to interact with technology. The system introduced a concept of Gesture Fusion, a novel approach that leverages the complementary strengths of Convolutional Neural Networks (CNN), Multi-Layer Perceptron (MLP), and the MediaPipe framework to achieve robust hand gesture recognition. By integrating these advanced techniques, fusion of models surpasses the limitations of individual model, delivering improved accuracy and robustness in recognizing a diverse range of hand gestures. The proposed system begins by preprocessing input images using CNNs to extract informative spatial features. Simultaneously, it utilizes the MediaPipe framework to accurately detect and localize key hand landmarks. These landmarks are then processed by MLPs to capture intricate patterns and variations in hand gestures that may not be fully captured by CNNs alone. By combining the outputs of both CNN and MLP models through a weighted fusion mechanism, synthesizes a comprehensive understanding of hand gestures, resulting in more precise and reliable recognition outcomes. Through significant research and evaluation on benchmark datasets, Fusion model outperforms independent CNN and MLP models. The findings obtained demonstrates the improved accuracy, resilience, and adaptability in a variety of real-world circumstances, confirming its potential to revolutionize the field of hand gesture detection. It advances the pursuit for more natural and intuitive human-computer interaction paradigms by seamlessly combining cutting-edge technologies and approaches.

Read the paper · More papers on PaperTik