BMformer: Enhancing 3D Human Pose Estimation with BiSSM and MS-TCN
Yunfeng Bai, Jinming Zhang, Huixian Qu · 2025
3D Human pose estimation is a challenging task because of its complex spatio-temporal dependencies and the structured nature of sequential data. Current methods often struggle with overlapping joints and rapidly changing poses. To overcome these challenges, we propose a novel model, the Bidirectional Multi-Scale Former (BMFormer), which incorporates the Bidirectional Structured Self-Attention Model (BiSSM) and the Multi-Scale Temporal Convolutional Network (MS-TCN). This method provides substantial advantages in modeling spatio-temporal correlations while capturing dynamics across varying temporal scales. The BiSSM model enhances the capacity to capture joint motion correlations through bidirectional temporal modeling combined with convolution-based feature extraction. At the same time, the MS-TCN module effectively captures short-term and long-term dynamics through a multibranch design, thereby enhancing the BMFormer's ability to process changes in motion on different temporal scales. BMFormer exhibits outstanding performance on two challenging benchmark datasets: Human3.6M and HumanEva. Particularly, our proposed model achieves state-of-the-art performance on the Human3.6M dataset using detected 2D poses from the Cascaded Pyramid Network.