Device-Free Gesture Recognition Using Multidimensional Feature Representation and Lightweight Self Attention-Free Transformer

Chao Fang, Yong Wang, Mu Zhou, Xiaolong Yang, Jiacheng Wang, Bao Peng · IEEE Transactions on Consumer Electronics · 2025

Device-free gesture recognition (DFGR) has always been a popular research topic in human computer interaction. Effective extraction and integration of multidimensional wireless signal features is crucial for fine-grained gesture recognition. Recently, deep learning (DL)-based DFGR methods, especially those using transformer networks, have achieved significant success in extracting global contextual features. However, DL-based DFGR methods for multidimensional data remain to suffer from various technical challenges, such as informative input features, lightweight network design, discriminative feature extraction and efficient fusion strategies for different multidimensional data fusion. To tackle these challenges, this paper proposes a novel end-to-end DL-baed DFGR fusion network, namely, multi-domain time-distributed convolutional neural network (TD-CNN) and lightweight attention-free transformer fusion network (MDFNet). With end-to-end training, the proposed fusion network effectively captures discriminative information from multiple domains, thereby enhancing the performance of joint recognition. Specifically, a learnable preprocessing module is designed to obtain informative multidimensional data, called RDA and RDE cube sequences. Subsequently, MDFNet takes a TD-CNN module to extract the range-Doppler-angel shallow features of RDA cube sequence and the range-Doppler-elevation shallow features of RDE cube sequence. The extracted features are fed into the well-designed encoder consisting of dilated convolutional module and interactive attention module derived from the conventional transformer encoder (DIFormer), which is crucial in fusing multidimensional discriminative information within the network. Comprehensive experimental results demonstrate that MDFNet can achieve human gesture recognition F1-score of 99.72% and human activity recognition F1-score of 97.78%, which outperforms other comparative approaches.

Read the paper · More papers on PaperTik