Attention-based feature enhancement for direct multiview 3D human pose estimation
Peiling Song, Xuan Zhu, Xingwang Zhao, Jingjing Lei, Jiahao Zhu, Qian Dang, Lin Wang · Journal of Electronic Imaging · 2025
The task of 3D human pose estimation (3D HPE) is to estimate the coordinates of human joints from images or videos and connect adjacent joints to form a human skeleton. 3D HPE technology is widely used in the fields of behavior recognition, human–computer interaction, and virtual reality. The method of 3D HPE is divided into single-view method and the multiview method. The single-view method has great limitations in solving the problem of multiperson pose estimation and occlusion. The multiview 3D HPE can be classified into single-stage and mainstream two-stage methods. Compared with single-stage methods, the two-stage methods first estimate the coordinates of 2D joints in each view and then use the association between 2D joints lifting the 2D pose to the 3D pose, with an accuracy affected by the 2D estimation results. Moreover, the existing single-stage methods cannot consider the regression and attribution features of joints from a global perspective at the same time, with a performance that needs to be further improved. To address these problems, we propose an end-to-end multiview multiperson 3D human pose estimation network, named AFEMVPose. It uses attention mechanisms to adaptively focus on and enhance regression and affiliation features of joints, capturing complex interactions between joints. AFEMVPose consists of a feature extraction module (FE), attention-based feature enhancement module (AFE), and pose decoding module (DCPose). FE is used to extract the initial features of multiple views. AFE strengthens the regression features of joints to raise the localization precision of joints. The purpose of the specially designed DCPose is to enhance and integrate joint affiliation and regression features, achieving the correct connections of joints. Compared with the state-of-the-art methods, the 3D HPE accuracy of our method is competitive on the CMU Panoptic, Shelf, and Campus datasets and demonstrates good robustness.