Efficient Plug-and-Play Module for Transformer-Based 3D Human Pose Estimation
Ya‐Ru Qiu, Jianfang Zhang, Guoxin Wu, Yuanyuan Sun · 2024
Video pose transformers (VPTs) have demonstrated remarkable performance in 3D human pose prediction. However, transformer-based architectures are often computationally intensive, leading to prolonged training times. To address this issue, we propose an efficient and plug-and-play module named the Adaptive Token Management (ATM), to prune and recover pose tokens, enabling high efficiency and accurate predictions within the VPTs framework. ATM consists of two key modules: the Dynamic Token Pruning (DTP) and the Lightweight Token Recovery (LTR). The DTP is presented for dynamic tokens pruning, where it determines the optimal number of representative tokens based on feature distribution and selects the representative tokens through a scoring mechanism. The LTR module, an optional lightweight module, is designed for tasks requiring the recovery of full-length tokens. It selects an appropriate full-length token as a learnable initialization token, utilizing attention mechanisms to restore spatio-temporal information. We conducted extensive experiments on the Human3.6M and MPI-INF-3DHP datasets, demonstrating that our approach can enhance the efficiency of mainstream VPTs while maintaining prediction accuracy. The results show that, when using MixSTE as the underlying framework on the Human3.6M dataset, ATM accelerated training time per epoch, reduced FLOPs by over 40%, and decreased prediction error by 0.2.