InLite-HRNet: efficient band convolution network for human pose estimation
Ziyang Zhang, Cui Xie, Junyu Dong, Xiaofeng Chang, Yiguo Wang · 2025
In the task of fast human pose estimation, to obtain a broader receptive field and more comprehensive pose information, thetraditional method is to use a large convolution kernel in a lightweight network to expand the receptive field. However, largeconvolution kernels bring high computational complexity and are not suitable for fast estimation tasks. To address these problems, thisstudy presents a lightweight human pose estimation network named InLite-HRNet. It uses two orthogonal band kernels in the vertical and horizontal directions instead of a single large convolution kernel, effectively increasing the size of the receptive field without additional computation. This approach decomposes the depth-wise separable convolution of the large convolution kernel intofour parallel branches along channel dimensional (a square kernel, two orthogonal band kernels, and identity mapping) and embedsthem into a new module SyncroFuse to enhance its feature expression ability. SyncroFuse is the basic component unit of our InLite- HRNet that plays a significant advantage of using large Receptive field to extract features on the high-resolution network. Experimental results show that the proposed network achieves an average accuracy of 71.07% and 87.22% on COCOandMPII human pose estimation datasets, which is better than the existing excellent lightweight networks and comparable to the existing large- scale human pose estimation networks.