Enhancing 3D Human Pose Estimation Through Frequency-Spectral Adaptive Hierarchical Fusion Transformer

Hongda Li, Dandan Sun · 2025

Recent advancements in multi-frame enhancement methods have propelled the field of 3D human pose estimation. However, existing approaches often struggle with ambiguities arising from occlusion and noise in 2D pose sequences. To address these limitations, this paper proposes a Frequency-Spectral Adaptive Hierarchical Fusion Transformer(FSA-HFFormer). The proposed method introduces a hierarchical network for multi-level feature extraction and fusion across spatial, temporal, and frequency domains, incorporating the Frequency-Spectral Channel Adaptive Mechanism to effectively extract and integrate frequency domain features. This design improves motion representation for complex human poses. Additionally, a Spatio- Temporal-Frequency Domain Interaction Network is devised to further fuse these multi -domain information. Experimental results show FSA-HFFormer achieves MPJPE of 31.4mm (Human3.6M) and 50.2mm (MPI-INF-3DHP), with improvements of 4.6% and 13.4% over MHFormer, demonstrating significant advances in monocular 3D pose estimation.

Read the paper · More papers on PaperTik