Human Pose Estimation Combined with Transformer for Spatiotemporal Representation Learning

Feiyue Qiu, Delong Peng, Lin Sun, Jian Qiang Zhou · 2024

At present, the methods of dynamic human pose estimation have some problems, such as inadequate extraction of human pose features and inadequate capture of time information of pose features. To solve the above problems, we propose a human pose estimation method combined with Transformer for spatiotemporal feature representation. Firstly, multi-head attention module in Transformer is redesigned, and local and global multi-head attention relationship aggregation modules are designed in shallow layer and deep layer respectively. Secondly, in order to exchange information between different Windows and extract deeper features, a 3×3 deep convolution module is added to the Feed-Forward Network. Then, combined with high resolution network HRNet, dynamic position embedding (DPE) is introduced to ensure the flexibility of the input sequence. Finally, we obtain an average accuracy of 77.5% on the MS COCO 2017 pose estimation task, 97.7% and 76.7% on UCF101 and HMDB51 dynamic human pose estimation datasets, respectively. Experiments show that the proposed algorithm is superior to other algorithms in performance and its effectiveness is verified.

Read the paper · More papers on PaperTik