Enhanced Gesture Recognition Through Graph-Based Multimodal Fusion

Mobeen Ur Rehman, Talha Ilyas, Lakmal D. Seneviratne, Irfan Hussain · 2024

This study introduces an advanced framework for recognizing hand gestures from a first-person view, leveraging the integration of multimodal data including optical flow, pose, depth, and RGB video recordings. Adeptly navigating the challenges and opportunities presented by the integration of multimodal data. At its core, the framework employs two pivotal components: a cross-attention based adaptive graph convolutional network and relational graph interactions for modality fusion. The former is designed to extract features from skeleton-based gesture data, ensuring a nuanced capture of hand movements by emphasizing the interconnections within the hand's skeletal structure. The latter component innovatively models each output modality feature as a node in a fully connected relational graph, facilitating the fusion of heterogeneous data types through dynamic interactions between modalities. This approach allows for the leveraging of each data type's strengths and the mitigation of their weaknesses, significantly enhancing the system's classification accuracy and robustness. Tested on a public benchmark dataset, the framework achieved a remarkable accuracy of 98.48%, demonstrating its efficacy. Moreover, it proves resilient, maintaining strong performance (93.48% accuracy) even in scenarios where only one modality is available, highlighting its potential for real-world applications. This advancement sets a new benchmark in hand gesture recognition, promising future developments in multimodal data fusion.

Read the paper · More papers on PaperTik