Mixformer: 3D human pose estimation by combining frequency domain pose assumptions

Yuheng Wang, Chao Zhang, Tinglan Lu, Qi Li · 2024

In recent years, with the rapid development of neural networks, human pose estimation research has made remarkable progress. The ability of Transformers to establish element-level global dependencies in long sequences makes them particularly suitable for representing discrete human joints. However, due to the inherent ambiguity of 2D images and the oversight of multiple feasible solutions in previous pose estimation methods, ViT models face challenges in deployment on small- and medium-scale devices and often result in significant errors. To address these challenges, this paper proposes a novel Transformer-based pose estimation method, Mixformer. On one hand, Mixformer obtains compact representations of pose sequences in the frequency domain, ensuring sequence completeness while reducing computational overhead. On the other hand, the model generates multiple hypotheses to provide rich spatiotemporal information, thereby improving accuracy. Extensive experiments demonstrate that Mixformer achieves significant performance improvements on the Human3.6M dataset. In conclusion, the proposed Mixformer model efficiently and accurately performs human pose estimation, offering superior performance and broader applicability.

Read the paper · More papers on PaperTik