Local-Global Feature Fusion for Enhancing 3D Human Pose Estimation

Yuanhong Zhong, Guangxia Yang, Daidi Zhong, Xun Yang, Shanshan Wang, Zhangling Duan · IEEE Transactions on Circuits and Systems for Video Technology · 2025

Based on its excellent capability to extract temporal features, transformer has been widely used in monocular 3D human pose estimation. However, due to its global perspective, it performs inadequately in extracting spatial features, which hinders breakthroughs in performance. In this paper, we propose a local-global feature fusion method based on GCN and transformer for 3D human pose estimation. Our method integrates GCN with multiscale transformer to extract local spatiotemporal features of poses. These are then integrated with the global spatiotemporal features extracted by vanilla transformer to reconstruct 3D human poses accurately. In addition, we introduce a hierarchical feature fusion method to better capturing the underlying 3D pose structure. It blends deep abstract features with shallow raw features. We evaluate our model on the Human3.6M and MPI-INF-3DHP datasets, and experimental results demonstrate that our approach outperforms existing state-of-the-art methods. We achieve advanced performance on both datasets with errors of 37.7mm and 16.4mm under MPJPE, respectively. The code and model are available at https://github.com/ygx7/LG3DPose.

Read the paper · More papers on PaperTik