Recurrent Transformer for 3D Human Pose Estimation
Guang Cheng, Yan Huang, Bing Yu · 2023
Currently, common three-dimensional (3D) human pose estimation algorithms achieve good results in representation learning, but still suffer from poor estimation accuracy and depth ambiguity at the joint points of the human skeleton, and extracting the image context is highly promising for mitigating the depth ambiguity. Therefore, an effective way to estimate human pose from monocular video images using redundant two-dimensional (2D) pose sequence spatio-temporal information is a research challenge. Many existing studies have mostly attempted to capture the spatial as well as temporal relationships of human poses in videos to solve these two problems. However, these works tend to overlook the fact that the lack of remote dependency modeling capabilities and parallelism makes the estimated 3D human poses often less accurate. We have proposed a Recurrent Transformer that can achieve a good balance between its efficiency and model size while still maintaining its effectiveness. This method will decompose the task into three phases: (i) serial cyclic spatial relationship modeling of the pose; (ii) parallel cyclic temporal relationship modeling of the pose; and (iii) summarization of multiple features and synthesis of the final 3D pose after linear regression. With the three main processes mentioned above, the final predicted 3D poses are not only improved compared to traditional methods but also more accurate in terms of precision. After experimental validation, Recurrent Transformer achieves leading results on Human3.6M.