3D Human Pose Estimation based on fused Spatio-Temporal Attention and Graph Convolutional Networks
Yubo Xue, Gang Chen, Zhanbo Li, Cong Zhang · 2025
3D human pose estimation is a widely used technique, and most existing models employ Transformer-based approaches due to their excellent performance. However, these models usually emphasise global relationships between human joints, which may compromise their ability to effectively capture local dependencies. In this paper, we propose a 3D human pose estimation model that combines spatio-temporal attention with graph convolutional networks. The model contains two parallel dual-stream modules, Transformer and GCNFormer, which complement each other to enhance the model's capability for 3D human pose estimation through cross-fertilisation and adaptive fusion. Two variants, TGFormer and TGFormer-L, are proposed, which can be selected based on the desired trade-off between accuracy and speed. The efficacy of the model has been evaluated on two datasets, Human3.6M and MPI-INF-3DHP, with the TGFormer-L variant attaining state-of-the-art performance, exhibiting P1 errors of 39.5 mm and 16.3 mm, respectively. It is noteworthy that TGFormer attains high accuracy while maintaining a reduced parameter count, thereby signifying an optimal balance between accuracy and model complexity.