End-to-end 3D Human Pose Estimation with Transformer

Bowei Zhang, Peng Cui · 2022 26th International Conference on Pattern Recognition (ICPR) · 2022

Transformer based architectures have become the common choice in natural language processing and are now achieving SOTA performance in computer vision tasks such as image classification, object detection. However, the convolutional method still keeps SOTA performance in many approaches of 3D human pose estimation. Inspired by recent development in vision transformers, we design a heatmap-free structure using standard transformer architecture and learnable object queries to model the human joint relation within each frame and then output accurate joint positions and types, we also present a transformer based pose recognition architecture without any greedy algorithm to post-processing predicted bones during runtime. In the experiments, we achieve the best performance among methods that directly regress 3D joint position from a single RGB image, and report competitive results with many 2D to 3D Lifting approaches.

Read the paper · More papers on PaperTik