A Novel Auxiliary Task Framework in 3D Human Pose Estimation for Opera Videos

Xingquan Cai, Haoyu Zhang, Shanshan He, Haoyu Song, Haiyan Sun · 2024

Influenced by the costume and limb occlusion of dance movements in opera videos, the 2D human pose estimation methods struggle to accurately locate the 2D joint coordinates of the occluded parts. This inaccuracy leads to lower precision in estimating 3D human poses from 2D joint coordinates in opera videos. To enhance the learning of more effective spatio-temporal dependencies from 2D joint coordinates and improve the accuracy of 3D human pose estimation, this paper proposes a novel auxiliary task framework. We first designed three auxiliary tasks to mask some of the 2D joint coordinates, disorder the spatial position of the joints and the video frame order, with the goal of recovering the corrupted 2D joint coordinates. Then to address these auxiliary tasks, we propose a multi-feature representation Transformer network to capture 2D joint coordinates spatio-temporal features from local to global perspective by constructing local adaptive graph convolution network, segmented time-aware network and global spatio-temporal self-attention module respectively. Finally, an adaptive weight allocation module is utilized to integrate local and global features to output the 3D joint coordinates. Extensive comparative and ablation experiments on the Human3.6M, MPI-INF-3DHP and opera datasets demonstrate that our method surpasses all comparative methods in MPJPE accuracy. Furthermore, the auxiliary task framework designed in this paper effectively captures comprehensive and efficient spatio-temporal dependencies in 2D joint coordinates from opera videos.

Read the paper · More papers on PaperTik