ST-CFNet: Spatio-Temporal Cross-Feature Fusion Networks for 3D Human Pose Estimation
Shu Cao, Chong Cao, Wei Zhao · IEEE Signal Processing Letters · 2025
Despite the remarkable progress in 3D human pose estimation, existing methods face dual challenges of insufficient spatiotemporal modeling and inefficient cross-modal interaction in complex motion and occlusion scenarios. To address this, we propose a novel network named ST-CFNet, which leverages a dual-stream interaction framework to separately capture the spatial structural features and temporal dynamic patterns of the human body. Meanwhile, we design an Adaptive Cross-feature Fusion Module (ACFM), which constructs a decomposable channel interaction matrix to adaptively enhance spatiotemporal features and achieve cross-modal fusion, while suppressing irrelevant information. Compared to traditional parallel architectures, ST-CFNet enables progressive fusion of spatiotemporal cues during the hierarchical feature extraction process. Experimental results show that the model achieves significant performance improvements on benchmark datasets such as Human3.6M and MPI-INF-3DHP, while reducing estimation error.