FPNet: Fusion Attention Instance Segmentation Network Based On Pose Estimation
Lei Pi, Jin Wu · 2021
Instance segmentation is a very challenging task. The standard approach to image instance segmentation is based on object detection, such as Mask R-CNN. It can't be good at handling the occlusion between people. Moreover, the edge information is not rich enough, and the capture of useful information in the feature extraction process is not efficient. In order to solve these problems, we propose FPNet: a fusion attention instance segmentation network based on pose estimation. We add the Split-Attention module to the backbone, so in the feature extraction process, the focus position can be automatically selected according to our needs, and a more distinguishable feature representation can be generated, which can improve the performance of the entire network. In order to obtain richer edge information, we added the Point-based Rendering module to the segmentation module. The module can efficiently calculate high-resolution segmentation images, so the edge contours of the final output can be clearer. To verify the performance of the network, we tested it on COCOPerson(the person category of COCO) and OCHuman(Occluded Human). On COCOPerson validation set, FPNet reached 0.555 AP. On OCHuman validation set, our method reached 0.549 AP. On OCHuman test set, FPNet reacheded 0.544 AP.