RetinaMHSA: Improving in single-stage detector with self-attention

Sina Soleimani Fard, Abdollah Amirkhani, Mohammad Reza Mosavi · 2021

In recent years, object detection with two-stage methods is one of the highest accuracies, like faster R-CNN. One-stage methods which use a typical dense sampling of likely item situations may be speedier and more straightforward. However, it has not exceeded the two-stage detectors' accuracy. This study utilizes a Retina network with a backbone ResNet50 block with multi-head self-attention (MHSA) to enhance one-stage method issues, especially small objects. RetinaNet is an efficient and accurate network and uses a new loss function. We swapped c5 in the ResNet50 block with MHSA, while we also used the features of the Retina network. Furthermore, compared to the ResNet50 block, it contains fewer parameters. The results of our study on the Pascal VOC 2007 dataset revealed that the number 81.86 % mAP was obtained, indicating that our technique may achieve promising performance compared to several current two-stage approaches.

Read the paper · More papers on PaperTik