LiteAHNet: A Lightweight and Attention-Driven Network for Human Pose Estimation
Jiayu Zou, Shuren Zhou, Xinlan Duan · 2025
Human pose estimation is a crucial task in computer vision, which aims to accurately locate and detect key points of human body in images or videos. However, traditional high-resolution networks have difficulty in effectively capturing global context information during multi-scale feature integration, which limits the accuracy of feature point localization. At the same time, the high number of parameters and computational complexity also hinder its deployment on resource-constrained devices. To address these problems, our proposes the Lite Attention-High Resolution Network (LiteAHNet), which achieves model lightweighting and performance optimization by introducing a Dynamic Block (DyBlock) and an Attention-Inverted Residual MLP Block (A-IR). DyBlock reduces computational cost while maintaining efficient feature extraction, whereas A-IR enhances feature representation by integrating attention mechanisms with lightweight feature enhancement strategies. Furthermore, LiteAHNet employs a Multi-Scale Feature Fusion (MSFF) approach that combines transposed convolution and Laplacian of Gaussian filtering to improve the understanding and recognition of complex human poses. By reducing network modules and optimizing the architecture, LiteAHNet retains the advantages of high-resolution features while significantly lowering parameter count and computational complexity. Experimental results demonstrate its high practicality and efficiency on resource-constrained devices.