MHViTPose: multiscale hybrid vision transformer for human pose estimation
Junhui Qu, Ziyan Zhao, X. D. Yu, Wei Zhang · 2023
Despite the significant progress achieved by visual Transformers, there are still some limitations that need to be addressed in human pose estimation. Firstly, Transformer lacks CNN’s inductive bias and local feature attention capabilities, which require extensive training data and iterations to achieve satisfactory results. Therefore, we propose a hybrid network that combines convolutional and Transformer. Besides, to address the recognition of human body images at different scales, we established a Transformer pyramid structure, which achieves recognition of human body images at different scales through progressive reduction of the input resolution. Specifically, our algorithm achieves an accuracy of 77.3% with a computational complexity of 19.6 GFLOPs. Compared to traditional direct regression methods, our algorithm considerably enhances detection accuracy while reducing the training complexity and significantly increasing the detection speed compared to traditional Transformer methods.