Frequency-Enhanced Spatio-Temporal Criss-Cross Attention for Video-Based 3D Human Pose Estimation*
Xianfeng Cheng, Zhaojie Ju, Qing Hong Gao · 2024
Recent transformer-based solutions have achieved remarkable success in video-based 3D human pose estimation. However, the computational cost grows quadratically with the number of joints and frames when computing the joint-to-joint affinity matrix. In this work, we propose a novel Frequency-Enhanced Spatio-Temporal Criss-Cross Transformer model (FSTCFormer) to address the computational challenge from two perspectives: leveraging the compact frequency-domain representation of the input data and relevant learned decomposition. Our FSTCFormer first performs discrete cosine transform (DCT) and filtering on the input sequence to extract low-frequency coefficients, capturing trend information and successfully compressing the input data. Then, the Frequency-Enhanced Spatio-Temporal Criss-Cross (FSTC) block uni-formly divides the compressed features along the channel dimension into two partitions and separately performs spatial and temporal attention on each partition. A learnable Freq MLP is introduced during the attention computation to further enhance the utilization of frequency-domain data. Each FSTC block, by concatenating the outputs of the attention layers, can model the interactions among joints within the same frame, joints along the same trajectory, and joints fused across multiple frames. Extensive experiments on the Human3.6M demonstrate that our FSTCFormer achieves a better trade-off between speed and accuracy compared to state-of-the-art (SOTA) methods.