Vitcc: a Vision Transformer Coordinate Classification Perspective for Human Pose Estimation
Hui Liu, Qian Zheng · 2024
Simple coordinate classification (SimCC) utilizes coordinate classification to alleviate the long-standing quantization error problem in human pose estimation (HPE) but ignores the spatial relationships and prior knowledge hidden between coordinates. To tackle this issue, we propose Vision Transformer Coordinate Classification (VitCC), a more effective coordinate separation method. VitCC innovatively integrates the Vision Transformer (Vit) architecture into the HPE pipeline, transforming the keypoint detection task into a pixel input-based classification task. Specifically, this method converts keypoint heatmaps into different tokens and utilizes a Transformer with dynamic attention and global context performance for coordinate separation and prediction. Additionally, we introduce the Effective Sparse Modeling (ESM) transformer block to reduce semantic ambiguities in the self-attention mechanism and further model the local spatial context efficiently. Comprehensively experiments on two challenging datasets (COCO, MPII) demonstrate that VitCC outperforms current state-of-the-art coordinate classification methods.