SST: Simplified Space-Time Transformer Based on Time-Assisted Spatial MSA for 3D Human Pose Estimation
Sheng Lu, Qi-Yun Dong, Zhenyin Zhang, Gengsheng Chen, Yinna Zhu, Wei Xu · 2024
Depth ambiguity in 2D human joint estimation is a persistent issue for 2D-3D human pose estimation networks. To cope with this challenge, temporal dimensions are adopted in existing models, however, none of them are able to fully utilize the information embedded in the input data. In this paper, we present a Time-assisted Spatial (TaS) MSA & Simplified Space-Time Transformer (SST) to better capture the spatial-temporal relationships. First, we design a new Time-assisted Spatial (TaS) MSA to comprehensively model spatial-temporal relationships. Secondly, we combine TaS MSA and Temporal MSA in parallel to enhance modeling capability and to build Simplified Space-Time Transformer (SST) model. Thirdly, we find an optimal pipeline of SST through contrasting the impact of parallel blocks and intermediate feature dimensions on the model's performance. Experimental results show that our model achieves the highest accuracy on Human3.6M dataset, with 0.4mm gain against current methods and 9% improvement in difficult positions.