Multi-person Pose Estimation with Multi-Attention Mechanism
Qian Shao, Yeqin Shao · 2024
In recent years, multi-person pose estimation has emerged as a prominent research direction in the field of computer vision, holding significant importance in applications such as human-computer interaction, action analysis, and virtual reality. However, traditional methods are often complex and inefficient, particularly during the feature fusion process, which can lead to the loss of critical information and increased errors under occlusion and complex poses. To address this, multiple attention modules were introduced in the early stages of the network to enhance the modeling of dependencies between key points, ultimately overcoming the limitations of conventional heatmaps through a coordinate classification approach. The design utilizing multiple attention modules achieves a balance between maintaining a lightweight structure and improving accuracy. Furthermore, by introducing a multi-attention mechanism, information loss is reduced, thereby enhancing the model's robustness in handling occlusion and other challenges in complex scenarios. Compared to existing advanced methods, the approach presented in this paper achieves an average precision increase of 1.5 percentage points on the COCO dataset.