Combining GCN Attention and Frequency Domain Analysis for 3D Human Pose Estimation

Yue Li, Bo Zhang, Wendong Wang · 2024

The 3D human pose estimation algorithm has significant potential for use in the medical sector. Current research in this area primarily utilizes Vision Transformer architectures, which are generally organized into two phases: the extraction of spatial structural features from human keypoints and the analysis of temporal changes. In the past, researchers have used Graph Convolutional Networks (GCNs) to identify the relationships between various keypoints that make up the human skeletal structure or the same keypoint over time. However, a new challenge has arisen in integrating GCNs with the multi-head self-attention (MHSA) mechanism found in Vision Transformers to enhance the extraction of spatial and temporal dependency features. Additionally, existing research indicates that adding frequency domain information to Vision Transformers—particularly in time-related contexts—can lead to the extraction of more robust and comprehensive features that are less affected by noise. In light of this, we introduce the GFreFormer model, which utilizes a GCN-MHSA hybrid transformer block for the extraction of spatial and temporal features. It also incorporates traditional spatial transformer block and an enhanced Temporal-Fre Transformer block to better capture temporal features, ultimately achieving precise 3D human pose estimation results. Comprehensive experiments conducted on the Human3.6M dataset show that the GFreFormer model surpasses other existing methods.

Read the paper · More papers on PaperTik