3D Human Pose Estimation in Spatio-Temporal Based on Graph Convolutional Networks

Liming Jiao, Xin Huang, Lin Ma, Zhaoxin Ren, Wenzhuo Ding · 2024

In recent years, 3D human pose estimation has seen rapid development and plays a key role in fields such as human-computer interaction. Currently, the most widely used methods include monocular-based pose estimation and multi-camera-based pose estimation, with study on monocular methods being more extensive. Despite significant progress in 3D pose estimation from single-view images or videos, the accuracy and performance of 3D human pose estimation can be significantly reduced due to depth ambiguity and severe self-occlusion issues, making it still a challenging task. Considering that spatial dependency and temporal consistency can effectively alleviate these issues, this paper proposes a method based on graph structure to estimate 3D human poses from short sequences of 2D joint detection results. Specifically, the domain knowledge about human body structure is explicitly incorporated into the graph convolution operation to meet the specific needs of 3D pose estimation. The proposed method has been evaluated on the challenging benchmark dataset Human3.6M for 3D human pose estimation. Experimental results show that our method outperforms the traditional method using GCNs by 5%, achieving excellent results. This method shows potential applications in fields such as behavior recognition and human-computer interaction.

Read the paper · More papers on PaperTik