Multi-stage network based on connected simple backbone interaction for human pose estimation
Lingyi Cai, Wei Liu · 2021
In recent years, the task of human pose estimation has continuously made significant progress, and it has become one of the research hotspots in the field of computer vision. From a certain point of view, there are two types of human pose estimation methods: single-stage methods and multi-stage methods. Both of these two methods have good performance, but they also have shortcomings. Single-stage methods have encountered a common bottleneck that simply increasing the model capacity does not give rise to much improvement in performance. Because of the insufficiency in some design choices, the performance of multi-stage methods in current practice is not as robust as expected. Therefore, we balance the pros and cons and design a multi-stage human pose estimation network structure, each of which contains an excellent single-stage module. At the same time, we adopt strategies such as cross stage feature fusion and multi-scale supervision to effectively improve the accuracy of positioning key points of the human body. Finally, we applied this network to the COCO keypoint detection dataset, and proved the effectiveness of our network through analysis of related data.