3D Human Pose Estimation Transformer based on Spatio-Temporal Augmentation
Cong Zhang, Gang Chen, Yubo Xue · 2025
3D human pose estimation has been emerging in many cutting-edge fields, such as computer vision, autonomous driving, and robotics, which has triggered extensive attention in the industry and academia. In recent years, transformer-based methods have shown excellent 3D human pose estimation performance. The derivation of Q, K, and V vectors in their self-attention mechanisms in current 3D HPE methods are based on simple linear mappings. To address these issues, we incorporate an a priori attention module, which we model by constructing a spatial global topology and a temporal global topology. Based on this, we propose a new model, STGFormer, which first injects a priori knowledge into Q, K, and V via the STE module and then improves the perceptual ability of our model in the channel dimension and the spatial dimension via the CSA module. Then, the data is divided into two parts, and the temporal or spatial information is extracted from them. Finally, adaptive fusion is performed. By integrating these modules, our proposed architecture can effectively capture global and local information. Extensive experiments on two benchmarks (Human3.6M, MPI-INF-3DHP) show that STGFormer performs better than state-of-the-art approaches.