Optimizing Local-Global Dependencies for Accurate 3D Human Pose Estimation
Guangsheng Xu, Guoyi Zhang, Lejia Ye, Shuwei Gan, Xiaohu Zhang, Xia Yang · IEEE Transactions on Circuits and Systems for Video Technology · 2025
Transformer-based methods have recently achieved significant success in 3D human pose estimation, owing to their strong ability to model long-range dependencies. However, relying solely on the global attention mechanism is insufficient for capturing the fine-grained local details, which are crucial for accurate pose estimation. Existing local feature extraction networks, such as Graph Convolutional Networks, often suffer from over-smoothing, while small-kernel CNNs have limited receptive fields and are highly sensitive to 2D pose errors. These limitations constrain the full potential of data-driven approaches. To address this, we propose SSR-STF, a dual-stream model that effectively integrates local features with global dependencies to enhance 3D human pose estimation. Specifically, we introduce SSRFormer, a simple yet effective module that employs the skeleton selective refine attention (SSRA) mechanism, leveraging large kernels to capture fine-grained local dependencies in human pose sequences. This complements the global dependencies modeled by the Transformer, enabling a more comprehensive understanding of human motion. By adaptively fusing these two feature streams, SSR-STF can better learn the underlying structure of human poses, overcoming the limitations of traditional methods in local feature extraction. To the best of our knowledge, this is the first work to explore the application of large kernels in skeleton-based 3D human pose estimation. Extensive experiments on the Human3.6M and MPI-INF-3DHP datasets demonstrate that SSR-STF achieves state-of-the-art performance. Furthermore, the motion representations learned by our model prove effective in downstream tasks such as human mesh recovery. Codes are available at SSR-STF.