Global-Local Interplay with Transformer and GCN for 3D Human Pose Estimation

Dejia Xu, Tichao Wang, Fusheng Hao, Jun Sheng Cheng · Procedia Computer Science · 2025

The goal of 3D human pose estimation is to use pictures or videos to determine the 3D spatial locations of human joints. Transformer and GCN are the two key techniques to understand the relationship between the nodes in poses. However, methods based on transformers, which depend on self-attention for modeling global relationships, lack adequate consideration for adjacent local nodes. GCN-based methods are adept at local joint dependencies via skeletal adaptability, but fail to effectively capture global pose correlations. We suggest a Transformer-GCN fusion model with adaptive selection of spatial-temporal global-local features for 3D human pose estimation based on monocular videos in order to overcome these drawbacks. The model first constructs a parallel branch structure comprising Transformer and GCNFormer, efficiently integrating their respective strengths in capturing global semantics and local structures. Furthermore, each branch is equipped with two serial processing streams: the Spatial-Temporal Stream (ST-Stream) and the Temporal-Spatial Stream (TS-Stream). These dual-order streams work collaboratively to learn latent 3D human structural features through different processing sequences, thereby enhancing the model’s capability to model complex spatio-temporal correlations. Finally, we adopt learnable vectors to measure the weights of spatial-temporal global-local features, enabling adaptively selective extraction of these features. Experiments on the Human3.6M and MPI-INF-3DHP datasets demonstrate that our suggested approach produces state-of-the-art results, with MPJPE metrics of 37.7 and 16.7, respectively.

Read the paper · More papers on PaperTik