LAGD-Net: Second-Order Motion Representation Learning for RGB-D Motion Recognition
Benjia Zhou, Yiqing Huang, Yanyan Liang, Du Zhang, Jun Wan · 2024
Motion recognition presents challenges beyond image classification due to the need for spatiotemporal analysis. Traditional methods like 3D convolutions often fail to capture comprehensive video representations because of the complex space-time interactions and limited video data. While existing approaches decouple spatial and temporal domains using local (e.g., S3D, R(2+1)D) or global (e.g., UMDR-Net) strategies, their integration remains underexplored and primarily focuses on first-order motion features, leading to potential temporal misalignment. To address this, we introduce LAGD-Net, a novel architecture that combines Local short-term sequence Aggregation and Global domain Disentanglement for second-order motion representation learning. Our method integrates both decoupling strategies, using local 3D convolutions for spatial feature extraction and local temporal modeling (first-order), followed by spatiotemporal decoupling to capture global motion relationships (second-order). Experiments on action and gesture datasets demonstrate our method achieves state-of-the-art performance.