HR-MViT: Human Pose Estimation via Lightweight Attention Mechanism

Lijun Kou, Bingchen Li · 2025

Conventional human pose estimation models are often difficult to deploy on mobile devices due to their large number of parameters. Although CNNs (convolutional neural networks) offer strong performance, convolution operations are limited in capturing global and long-range semantic dependencies. To address these challenges, we propose we propose HR-MViT (High-Resolution Mobile Vision Transformer), a human pose estimation model that replaces traditional convolutional backbones with a lightweight Transformer architecture, MobileViT, within an HRNet-like framework. HR-MViT retains the core design philosophy of HRNet by adopting a multiresolution parallel structure and maintaining high-resolution representations throughout the network. For computational efficiency optimization, we employ MobileViT-incorporating depthwise separable convolutions-as the feature extraction backbone, while adaptively configuring module quantities across branches to achieve an optimal efficiency-performance equilibrium.We validate HR-MViT on the MPII (Max Planck Institute for Informatics) Human Pose dataset, and the results demonstrate that our model achieves significant lightweighting compared to HRNet while meeting real-time performance requirements. HR-MViT is suitable for deployment on mobile devices.

Read the paper · More papers on PaperTik