Multi-Person Estimation Method Combining Interactive and Contextual Information
Zhuxiang Chen, Shujie Li, Wei Jia · 2024
Multi-person pose estimation based on monocular cameras is one of the hot research topics in computer vision. Current monocular multi-person 3D pose estimation methods often treat individuals as independent entities for estimation and are based solely on single-frame assessments. This methodology has notable limitations, particularly its insufficient utilization of the rich interactive information interlinking individuals, as well as the contextual continuity inherent across sequential frames. To address these challenges, our study introduces a novel method for monocular multi-person 3D pose estimation. Firstly, it utilizes the attention mechanism of Transformer to learn the contextual and interactive information of both local and global coordinates. Subsequently, the methd employs low-order convolution operators of a root trajectory network to further assimilate global coordinate information. Finally, an optimization algorithm is used to adaptively filter frames with significant occlusions. Our method effectively integrates the interactive information among multiple individuals and the contextual information across multiple frames. Experimental results on the public dataset MuPoTS-3D demonstrate that compared to the state-of-the-art methods, our approach shows a 4.8% improvement in average joint accuracy and a 19. 4% reduction in average root joint error.