Real-Time Hand Gesture Recognition System Using Mediapipe, Spiking Convolution Neural Network and Bi-LSTM
Avinash Dhiran, Anurag Kumbhare, Achal Patil, Mrugank Vichare, Dhananjay Patel · 2025
World Health Organization claims that approximately 5 % of the world population suffers from being deaf and mute. This highlights the critical need for technologies to bridge the communication gap between individuals with speech and hearing impairments and the rest of the world. However, the development of sign language recognition systems is hindered by the scarcity of suitable datasets, as each spoken language has its own distinct sign language, along with the high computational power required for the implementation. This paper presents a novel real-time dynamic hand gesture recognition system that creates robust and customizable datasets, preprocesses the data and recognizes hand gestures, making it adaptable to any sign language. Currently based off the Indian Sign Language (ISL), the model uses MediaPipe and Spiking Convolution Neural Network (SCNN) for spatial feature extraction and a Bi-directional Long Short-Term Memory (Bi-LSTM) architecture for temporal features. The model is trained on 10 gestures with 100 diverse videos per gesture, that are 2 -seconds each. The proposed system achieves a high accuracy of 96 %, compared to the benchmark MediaPipe, ResNet50 and Bi-LSTM architecture, which attains 97.3 %. While the CNN and Bi-LSTM model provide slightly higher accuracy, the proposed system sheds new light on the potential of SCNNs in computer vision with frame-based inputs while also ensuring lower computational power consumption than state-of-the-art CNN models.