Transformers With Joint Tokens and Local-Global Attention for Efficient Human Pose Estimation
Kaleab A. Kinfu, René Vidal · IEEE Transactions on Pattern Analysis and Machine Intelligence · 2026
Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) have driven significant progress in 2D human pose estimation. However, achieving a good balance between accuracy, efficiency, and robustness remains challenging. For instance, CNNs are computationally efficient but struggle with long-range dependencies, while ViTs excel in capturing such dependencies but suffer from quadratic computational complexity with respect to the number of patch tokens. This paper proposes two ViT-based models for accurate, efficient, and robust 2D pose estimation. The first, EViTPose, improves computational efficiency with minimal accuracy loss by utilizing learnable joint tokens to select and process a subset of patches most relevant to the body joints, thereby enabling control over the accuracy efficiency trade-off by varying the number of processed patches. The second, UniTransPose, while not providing the same direct control over this trade-off, efficiently handles multiple scales by combining (1) an efficient multi-scale transformer encoder using both local and global attention and (2) an efficient sub-pixel CNN decoder for improved speed and accuracy. Furthermore, by incorporating joint annotation schemas from different benchmarks into a unified skeletal representation, we train robust models that learn from multiple datasets simultaneously and perform well across a range of scenarios, including variations in pose, lighting, and occlusion. Experiments on six benchmarks demonstrate that the proposed methods substantially outperform state-ofthe-art methods while improving computational efficiency. EViTPose significantly reduces computational cost (reduces GFLOPs by 30%44%) with minimal drop in accuracy (0% to 3.5%), and UniTransPose achieves accuracy improvements ranging from 0.9% to 43.8% across these benchmarks.