Hierarchical Contrastive Consistency for Human Pose Estimation in Images and Videos

Xixia Xu, Qi Zou, Jiamao Li · IEEE Transactions on Circuits and Systems for Video Technology · 2025

Human pose estimation (HPE) is an invaluable task in computer vision with various practical applications. This paper proposes a novel Hierarchical Contrastive Consistensy constraint (HICCON) to improve the HPE in both images and videos, which describes the input into multi-granular representations at spatial and temporal domain and performs multi-level feature consistency by exploring the characteristic of human structure and time sequence. The hierarchical contrast is conducted at four levels: keypoint-level, part-level, instance-level and clip-level. In spatial, we consider keypoint-level and part-level consistency across instances within frame to enhance the fine-grained keypoint robustness. The former conducts the single keypoint feature contrast across instances to improve the category-specific keypoint features. The latter explores the specific pair-wise features for preserving the instructive relation. In temporal, we develop the instance-level and clip-level feature consistency across frames to capture more discriminative temporal representations. The former discriminates instance features across frames within the same video, whereas the clip-level constraint aims to discriminate consistent features from different videos in order to capture more distinctive temporal features. Extensive experiments on kinds of architectures across datasets i.e, PoseTrack2017, PoseTrack2018 and PoseTrack2021 show the HICCON achieves about 1.5% improvement than baseline. Besides, the proposed method unleashes the potential of the contrastive learning in HPE field.

Read the paper · More papers on PaperTik