Advancing Three-Dimensional Human Pose Estimation Through Spatiotemporal Feature Fusion in Graph Convolutional Networks
Luoyang Chen, Junxian Li, Ye Tao, Yazhe Cheng, Wenchao Du, Zheng Liu, Xingdong Bao, Hongxia Mao · Advances in transdisciplinary engineering · 2025
Accurate three-dimensional (3D) human body pose estimation from video imagery is critical for a variety of applications, including action recognition, body language interpretation, motor skill acquisition, and motion capture. Despite considerable progress in this area, current methodologies often struggle to effectively integrate spatiotemporal features, resulting in limitations in both pose estimation accuracy and computational efficiency. Addressing this gap, we propose a novel graph convolutional network that synergistically fuses spatiotemporal properties to enhance 3D human pose estimation. By utilizing a sequence of human 2D poses as input, we extract spatial features through a semantic map convolutional network and temporal features using a time domain convolutional network. Additionally, we introduce a grouping Top K pooling method to optimize the extraction of multi-scale structural features, significantly reducing model parameters while enhancing pose estimation accuracy. Experimental evaluations on publicly available datasets demonstrate that our approach achieves highly accurate 3D pose estimation with real-time processing capabilities. This research not only provides a robust solution 3D pose estimation but also advances the field by improving the integration of spatiotemporal features, thus enhancing applicability in real-world scenarios.