Utilizing Spatial Transformers and GRU for Temporal Context in 3D Human Pose Estimation

Cheng Chen, Huahu Xu, Jian Kang · 2024

This paper presents a novel approach for 3D human pose estimation by integrating spatial multi-scale features with GRU-based temporal multi-scale encoding. Our method addresses the limitations of existing single-stage and multi-stage approaches by enhancing feature extraction and reducing redundancy. We replace the HRNet fusion layer with a multi-resolution module that generates multi-scale features through parallel branches. Temporal multi-scale encoding further refines the 3D poses by adaptively selecting information based on relationships between different resolutions. Experimental results on the Human3.6M and COCO datasets show our approach outperforms state-of-the-art methods, achieving 43.6 mm and 35.4 mm MPJPE on Protocol #1 and Protocol #2 respectively, and an AP of 75.2 on the COCO val 2017 dataset. These results demonstrate significant improvements in pose estimation accuracy and robustness.

Read the paper · More papers on PaperTik