Multiview Human Pose Estimation via Temporal Fusion With Attention Adaptive Fusion Weights
Yang Gao, Shigang Wang · IEEE Sensors Journal · 2025
Human pose estimation remains a critical and actively evolving research area in computer vision. While multi-view pose estimation has shown strong potential in addressing occlusion-related ambiguities, most existing approaches rely on fixed-weight fusion strategies that fail to adapt to varying view qualities. This often leads to suboptimal joint representations, particularly in the presence of occlusion or noise. To address this issue, we propose a novel multi-view human pose estimation approach that integrates multi-temporal fusion with attention-based multi-view fusion. Specifically, a multi-frame temporal fusion module that exploits motion context from consecutive frames and refines heatmaps for more robust keypoint detection. An attention-guided multi-view fusion mechanism that dynamically adjusts view weights based on joint confidence scores, enhancing cross-view feature alignment. An epipolar geometry-constrained correspondence module that leverages the sparsity of heatmaps to achieve more accurate keypoint matching. Extensive experiments demonstrate that the proposed approach outperforms several popular or state-of-the-art approaches in both quantitative metrics and visual quality across multiple public datasets. In particular, it achieves a mean per joint position error (MPJPE) of 9.17mm on the Occlusion-Person dataset, setting a new benchmark for pose estimation under occlusion.