Cross-Modal Vision-Language Model for High-Precision Human Pose Estimation
Hai Liu, Chengyue Bai, Tingting Liu, Yu Song, Li Zhao, Hong Chen, Linli Tan · 2025
Human pose estimation (HPE) is a critical task in computer vision, with applications in behavior monitoring, human-computer interaction, action recognition, and online learning. However, practical HPE applications often face challenges such as object occlusion, crossed limbs, blurred backgrounds, and complex backgrounds, which significantly impact the accuracy of keypoint localization. Traditional handcraft feature-based approaches and convolutional neural network (CNN)-based methods have limitations in handling these challenges. This paper proposes a novel Cross-Modal High-Precision Human Pose Estimation (CMP-HPE) method that leverages a cross-modal transformer to integrate visual and language information, enhancing the robustness and accuracy of HPE. The proposed model comprises four main modules: visual tokens construction, language tokens construction, cross-modal transformer, and marker-based prediction. Experiments on the MSCOCO and MPII datasets demonstrate the superior performance of our method, achieving significant improvements over state-of-the-art techniques. The source Python code will be available upon request.