HR-xNet: A Novel High-Resolution Network for Human Pose Estimation with Low Resource Consumption
Cun Feng, Rong Zhang, Lijun Guo · 2024
In our study, we are aiming to find an effective and lightweight solution for the human pose estimation task. To this end, a novel high-resolution representation network called as HR-xNet is proposed. First, HR-xNet is derived from the HRNet [20], by replacing the stem and Stage 1 of HRNet with the lightweight feature extraction head and pruning HRNet. Then, our focus is on improving the model's representation ability for changeable human poses and small targets such as human keypoints, while continuing to reduce model's resource consumption. First, the novel Lightweight Multi-Scale Dynamic Convolution (LMSD Conv) is introduced. The LMSD Conv greatly improves the learning capacity of the network by adaptively generating convolution kernels of different sizes to extract features with different receptive fields from the input. Second, low-level detailed and high-level semantic features are interacted with in the Feature Enhancement Module to relearn the lost detailed features for small keypoints and changeable human poses. On the COCO and CrowdPose datasets, our model can compete with some mainstream large networks and existing state-of-the-art lightweight methods at a low cost.