3D Human Pose Estimation with a Spatio-Temporal Cross Attention-GCN Network
Yucheng Xiong, Qishen Li, Qiufeng Li, Hua Huang · 2025
This study focuses on 3D human pose estimation in order to enhance estimation accuracy by integrating graph convolution and self-attention mechanisms. 3D human pose estimation has critical applications in fields such as intelligent surveillance, motion capture, and human-computer interaction, and involves extracting human joint features and constructing skeletal models from video data. This study proposes a Spatio-Temporal Cross Graph Convolution (STCG) block that applies graph convolution across both spatial and temporal dimensions, significantly improving the model's efficiency and its capacity to model spatiotemporal relationships. Furthermore, this study proposes a dual-stream architecture that combines STCG with a Spatio-Temporal Cirss-cross attention (STC) block, which results in a Spatio-temporal Cirss-cross attention and Graph Convolution (STCAG) block. The model leverages both graph convolution and attention mechanisms to achieve a comprehensive Spatio-Temporal representation. The nested dual-stream structure effectively enhances the model's understanding of human skeletal structures and its ability to capture motion trajectories. Experimental results on the Human3.6M dataset demonstrate that the STCAG framework shows superior adaptability in complex scenarios and reduces estimation errors.