A Novel 3D Human Pose Estimation Model Integrating Graph Convolution Networks and Attention Mechanism
XinFeng Chen, Wei Long, EnTuo Liu, Lingxi Hu, Attila Vidacs, LinHua Jiang · 2025
Monocular 3D human pose estimation (3D HPE) has emerged as a significant area of research within computer vision. The recent integration of Transformer architectures has opened new avenues for advancing this field; however, challenges such as potential occlusion and inherent depth ambiguity remain prevalent. This paper introduces a novel 3D human pose estimation network that synergizes the attention mechanisms of Transformers with graph convolutional networks. By capitalizing on the interconnectivity of joints, we implement structure-dependent modeling, which enhances the extraction of limb features. Additionally, our approach incorporates both graph convolution and attention mechanisms to optimize feature extraction efficiency. We employ a parallel structure to bolster the capture of contextual information and mitigate the accumulation of joint estimation errors. Experimental evaluations on the Human3.6M dataset demonstrate that our method achieves improvements of 0.7% and 6.1% in two critical evaluation metrics during training and validation, respectively, without the need for fine-tuning compared to existing approaches. These findings underscore the effectiveness and robustness of our proposed method, highlighting its potential for real-world applications in 3D human pose estimation.