ConFormer: Convolutional Transformer Exploiting Spatial and Temporal Information for 3D Human Pose Estimation

Hongde Luo, Nian Gu, Heng Zhang · 2022

Transformer has recently been introduced into 3D human pose estimation by exploiting spatial and temporal information. Though Transformer-based methods have made some considerable progress, they still leave some potential improvements on the table, such as poor locality, ignoring channel adaptability, and computational overhead. In this paper, we propose a spatial-temporal convolutional Transformer architecture, which lifts a long continuous 2D pose sequence to one 3D pose. Specifically, a spatial interaction module is adopted to model spatial relations between body joints within each frame and a temporal interaction module is applied to capture temporal dependencies across frames. Besides, we apply a special kernel attention mechanism in the spatial interaction module to explicitly encode the local and global relations between the body joints. In the temporal domain, the convolutional position embedding and the convolutional projection are utilized to enhance the ability to model local relationships and decrease computational complexity. The evaluation experiments of our method are conducted on two popular 3D human pose estimation datasets, Human3.6M and MPI-INF-3DHP. Results show that our proposed approach achieves better performance.

Read the paper · More papers on PaperTik