Human Pose Estimation Network Based On MultiHeatmap Fusion

Yangguang Zhao, Jun Lu · 2024

Estimating human pose in complex multi-frame situations is a challenging task and has attracted intensive research by many researchers. Although 3D human pose estimation methods have achieved remarkable results in scenes based on single images, their performance often fails once these model transformations are applied to video sequences. Common problems with these models include inability to cope with motion blur, out-of-focus videos, and occlusion of human poses. In order to solve the above problems, this paper proposes a feature extraction and representation model MHMF, which is used in the feature extraction stage of the model. The initial features extracted by the backbone network HRNet-w32 are guided by the heat map to the attention layer, which improves the network’s attention to important areas for predicting key points of human posture. At the same time, the integration of the aggregation heat map and the backbone network heat map improves the spatiotemporal consistency of the key points. In addition, to improve the accuracy of mesh pose estimation under occlusion, this paper proposes a transformer-based NewDSTformer model. By adjusting the structure of the Transformer encoder, increasing the encoder level and combining it with the dynamic progressive attention masking method. The model can adapt to different input situations, handle the positional relationship of local key points, and be able to perform accurate detection even under occlusion. It was evaluated on the 3DPW data set and improved the accuracy by $0.3 \%$, indicating that this paper effectively improved the performance of 3D human mesh reconstruction.

Read the paper · More papers on PaperTik