STAPFormer: A New 3D Human Pose Estimation Framework in Sports and Health
Zhongteng Zhang, Qing Peng, Liu Zhang, Zihao Zhang, Weihong Huang · 2024
3D human pose estimation is a fundamental technology for capturing human motion, providing precise skeletal position information and essential support for fields such as sports analysis, medical rehabilitation, and virtual reality. Recent advances in transformer-based models have demonstrated remarkable success in the domain of 3D human pose estimation. However, existing methods tend to focus predominantly on leveraging spatio-temporal correlations by capturing the inter-frame relationships of individual joints and the identical joint moving across frames, which may not fully harness the complex interactions between different joints across multiple frames. To alleviate this limitation, we introduce a novel SpatioTemporal Aggregation (STA) block designed to capture features based on Spatio-Temporal Aggregation of Pose (STAP). This block adeptly aggregates contextual information from both spatially and temporally proximate joints, effectively modeling the intricate relationships across the entire pose sequence. Technically, the STA block operates by directing the input feature through the two parallel transformation streams, which are tailored to capture the joint-based and STAP-based interactions. The outputs from these transformation layers are then integrated to pay more attention on the dynamic movement patterns of body parts and the overall pose sequence, yielding superior predictive accuracy in 3D pose estimation task. We propose STAPFormer, a model architecture that employs multiple stacked STA blocks, and conduct evaluations on established benchmark datasets. Our model achieves competitive results on the Human3.6M dataset and state-of-the-art performance on the MPI-INF-3DHP dataset, matching existing state-of-the-art methods. In addition, we conduct an ablation study that integrates various configurations of the joint-mixer with the proposed STA block. This study validates the effectiveness of proposed STA block in enhancing 3D human pose estimation. Code and models will soon be available at https://github.com/z2tng/stapformer.