A Robust Framework for 3D Human Pose Estimation Using Semantic Graph Convolution, Criss-Cross Attention and Transformer Encoder
Yu Hou, Cenyue Wang, He Peng, Tong Feng, Haoxuan Li, Youn Poong Oh · Journal of Circuits Systems and Computers · 2025
Accurate 3D human pose estimation is critical in various applications, such as motion capture, human–computer interaction and sports analysis. However, existing methods often struggle with capturing complex joint dependencies, especially in scenarios involving occlusions or diverse poses. To address these challenges, we propose GCAT-Net, a novel framework that integrates Semantic Graph Convolutional Network (SemGCN), Criss-Cross Attention (CCA) and Transformer Encoder to effectively capture both local and global dependencies in human joint relationships. GCAT-Net’s ability to combine these components allows it to accurately predict 3D poses, even in complex environments, making it suitable for a wide range of real-world applications. Experimental results demonstrate that GCAT-Net outperforms existing methods on multiple benchmark datasets. The model exhibits strong generalizability, handling complex poses and occlusions effectively, which highlights its robustness and adaptability. GCAT-Net not only improves estimation accuracy but also contributes to the understanding of human motion by enhancing feature extraction at both local and global levels.