LCGPose: Local-Cross-Global Transformer for 2D Human Pose Estimation
Jun Yin, Xiong Tao, Huadong Pan · 2022 IEEE 2nd International Conference on Data Science and Computer Application (ICDSCA) · 2022
2D Human pose estimation aims to locate the human keypoints from input data, and most CNN-based methods have achieved good performance for human pose estimation, however, the CNN-based methods lack the ability of capturing the relationship between different human keypoints, which limits the model performance to some extent. In this paper, we combine the beneficial qualities of CNN in processing the low-level vision and the advantages of Transformer in handling the relationship between various vision elements or objects, and then propose a new model framework-LCGPose for 2D human pose estimation. In detail, LCGPose mainly consist of Local Transformer Module, Cross Transformer Module and Global Transformer Module, which are respectively adopted to capture local dependencies, cross dependencies and global dependencies by self-attention mechanism of Transformer. A large number of experiment results show that our proposed LCGPose achieves competitive performance compared with the state-of-the-art methods via fewer model parameters and GFLOPs.