Cross-Modality Gesture Recognition With Complete Representation Projection

Xiaokai Liu, Mingyue Li, Boyi Zhang, Luyuan Hao, Xiaorui Ma, Jie Wang · IEEE Internet of Things Journal · 2024

Human gesture recognition, due to its indispensable role in a myriad of emerging applications, has attracted close attention in the visual and wireless sensing community. The intrinsic characteristics of visual and wireless modalities are complementary to each other, e.g., wireless signals are robust to illumination changes and occluded conditions but suffer from low-spatial resolution, while visual signals have a high-spatial resolution but are vulnerable to scenario variations. Intuitively, integrating them has the potential chance to improve the overall discriminative ability. However, due to their different physical patterns and semantics, how to explore the relationship between the two modalities and leverage their complementary information to improve recognition performance still remains unsolved. In this article, we propose a complete representation projection method, which projects the signals from the heterogeneous modalities into a complete representation feature space by performing a bidirectional projection constraint. Furthermore, to relieve the low-projection efficiency problem caused by the heterogeneity of the two modality features, we propose an attention-based cross-modality interaction (ACMI) mechanism to perform implicit semantic feature alignment, thereby better capturing the complex dependencies between the two modalities and improving the feature projection efficiency. To evaluate the proposed method, we build a visual-radar cross-modality gesture recognition system and conduct extensive experiments. Experimental results demonstrate that the proposed approach not only performs favorably against vision-only and wireless-only solutions by a large margin, but also shows superiority over traditional fusion solutions.

Read the paper · More papers on PaperTik