sEMG-vision Tra: A Gesture Recognition Method Based on Surface EMG Signal-Vision Fusion

Weili Ding, Shuhui Zhang, Han Liu · 2024

To address the challenges of gesture recognition accuracy within human-computer interaction technologies, and to overcome the limitations of single-modal biometric feature recognition - which often falls short in practical applications, this paper introduces a multimodal feature fusion classification algorithm, sEMG-vision Tra. This algorithm aims to integrate different types of biometric data, thereby enhancing the robustness and effectiveness of gesture recognition. Initially, we utilize the Mediapipe framework to extract skeletal point data from images of the human hand. The next process involves calculating distance and angle features between skeletal points, as well as extracting time-domain features from surface electromyographic (sEMG) signals. Following this, features from these two modalities are concatenated along the feature dimension and then linearly projected onto the target dimension. To effectively capture the intricate interrelationships among these features, we employ a multi-layer Transformer encoder for encoding the combined features. The encoded features are subsequently mapped to the final output dimension. The effectiveness of our approach is demonstrated through validation on a custom dataset. This algorithm's key contribution lies in the fusion of two distinct modal features, images and EMG signals, to achieve a more holistic and precise feature representation. By this, it significantly enhances our ability to capture the essential information for applications like gesture recognition or human-computer interaction. The algorithm showcases exceptional performance, achieving a noteworthy accuracy rate of 95.56% on a custom dataset.

Read the paper · More papers on PaperTik