Enhanced Pose Recognition: Integrating Efficient Attention and Convolution in PoseFormer
Xin Yuan, Jinming Zhang · 2025
Recently, Transformer-based approaches have made notable advancements in the task of human pose estimation, particularly in transitioning from 2D to 3D. PoseFormerV2, as an innovative approach, substantially increased the receptive field and enhanced robustness. This approach makes only slight changes to PoseFormer, combining features from both the temporal and frequency domains, and successfully improving the speed and accuracy. However, in practical applications, the Attention module used in PoseFormerV2 is based on the dot-product attention mechanism. which integrates Q(query) and K(key), followed by processing with values (V). This approach struggles with handling complex input relationships, limiting the model's performance. To overcome this limitation, the study introduces an additive attention framework, substituting the conventional quadratic matrix multiplication and linear element-wise multiplication operations. Additive attention captures complex patterns between inputs by nonlinearly combining queries and keys, and it removes key-value interactions without sacrificing performance. This mechanism efficiently captures the interaction between queries and keys by utilizing combined linear projection layers, which are adequate for learning the relationships among input features, thereby improving the model's learning ability and expressive power. Building on this foundation, we also incorporate partial convolution operations to refine feature extraction and amplify the smoothing effect, leading to an overall improvement in performance. By combining the additive attention mechanism with partial convolution, we present the “PoseFormer-EAC” model, which sets new benchmarks across various performance metrics, particularly excelling in edge-node recognition, where it produces more precise results. Our experimental results on the Human3.6M benchmark dataset show that the performance of our method substantially exceeds PoseFormerV2 and other transformer-based models.