STPointNet for Human Action Recognition in MmWave Point Clouds
Chenliang Zhu, Jie Yang, Peiwei Deng, Junzhe Lin, Lianfen Huang, Hezhi Lin · 2024
It’s essential to effectively learn the spatial-temporal information of point cloud sequences for 3D action recognition, especially for the increasing popular application of Human Action Recognition (HAR) at smart home. In this paper, we utilize Frequency Modulated Continuous Wave (FMCW) radar to acquire point clouds and design the end-to-end network which can learn the spatial-temporal representation of dynamic 3D point cloud sequences, dubbed Spatial and Temporal Point Network (STPointNet). STPointNet is a permutation-invariant network that can learn from unstructured point clouds with irregular domains. The proposed STPointNet network utilizes the Spatial Part to extract 1024-dimensional global spatial features of point clouds and the Temporal Part to extract 1024-dimensional global temporal features of point clouds. These features are then concatenated to form a complete feature vector of 2048 dimensions, which is subsequently fed into a Multi-Layer Perceptron (MLP) with a non-linear activation function softmax to obtain classification scores. Extensive experiments show that the proposed method significantly outperforms existing state-of-the-art approaches according to the recognition accuracy, achieving a recognition accuracy of 98.26% on the MMAction dataset. Furthermore, it demonstrates excellence and robustness in utilizing sparse point clouds for 3D action recognition.