Human Parsing Network Algorithm Based on Improved YOLOv7

Zhiqiang Bao, Defei Du, Tong Tong, Siwei Wang · 2024

Aiming at the challenge of extracting human image features in single-stage instance level human parsing, which makes it difficult to perform detailed human parsing, a human parsing network based on YOLOv7 improved attention fusion contextual features is proposed to improve the accuracy of human instance parsing. An improved lightweight up-sampling operator was used in the neck network to provide a larger receptive field during up-sampling, enhancing the ability of feature maps to express human position information; In the feature fusion stage, an attention fusion context information module was proposed, which calculates the triple attention (T-Attention) of fine-grained feature maps and activates them through Pyramidical fusion-activate context (PFAC) to improve the semantic expression ability of feature maps. The experimental results show that this method can effectively improve the accuracy of human instance parsing. On the selected ATR (Active Template Regression) dataset, the parsing accuracy of each part of the human body reaches 50.2%, which is 2.8 percentage points higher than the improved YOLOv7 network mAP:0.5; The accuracy of human body parts parsing on the multi person dataset PPP (PASCAL Person Part) increased by 3.2 percentage points with mAP:0.5.

Read the paper · More papers on PaperTik