FANpose: 2D human pose estimation with fully attentional networks under vision transformer baselines

Mingliang Chen, Guangxing Tan · 2024

2D human pose estimation (HPE) has been a research focus of computer vision, 2D HPE baseline is one of the main studies. As the field of HPE continues to evolve, Vision Transformer Baselines have emerged as a significant area of interest, showing considerable potential in visual applications. However, accurate estimation remains a challenge in 2D HPE. This study introduces a novel approach named FANpose for 2D HPE in images. Building upon the top-tier VITpose Baselines, we innovate in two main aspects. Firstly, we employ fully attentional net- works to replace the vision transformer baseline model, thereby enhancing the model’s robustness. Secondly, we improve keypoint localization accuracy by replacing traditional Gaussian kernels with Laplacian kernels, thereby enhancing the model’s recognition precision. On the MS COCO dataset, our model achieves AP and AR scores that are respectively 0.4 and 0.6 higher than VITpose-B, and our model is 32M smaller in terms of parameters than VITpose-B.FANpose achieves satisfactory results in human pose estimation tasks, showcasing its immense potential for practical applications.

Read the paper · More papers on PaperTik