3D Human Pose Estimation with Two-Step MixedGraph Convolution Transformer Encoder

Lianfeng Hu, Jiwei Hu · 2025

In 3D human pose estimation, human joints typically exhibit continuity and interdependence during motion, suggesting that joint velocities and their positional relationships can offer valuable insights. However, the application of this information in 2D-to-3D pose estimation remains underdeveloped. To address this issue, we propose the TMGCTE (Two-step Mixed Graph Convolution Transformer Encoder) training design network, a method that integrates a variant of graph convolutional networks with a ransformer architecture. The input of TMGCTE is generated by fusing 3D velocity vectors and 3D zero vectors with 2D joint positions, creating two alternating 5D feature sets that are input into different 3D pose estimation modules. This allows for the effective use of 3D velocity vectors and 2D joint graph structure information during training, facilitating better learning of Spatio-Temporal features in the shallow layers of joints. Extensive experiments demonstrate that the models trained with the proposed TMGCTE framework significantly improve performance on the Human3.6M and MPI-INF-3DHP datasets.

Read the paper · More papers on PaperTik