FIoU loss: Towards Accurate One-stage Object Detector with Better Bounding Box Regression

Hao Zhao, Hao Deng, Jixiang Huang, Ju Zhou · 2024

Object detection is used to predict categories and locations of potential instances, which provides important decision-support for robots. Compared with the classification sub-task, it’s more difficult to obtain accurate location in the regression sub-task because of two drawbacks. Firstly, the backbone of detector is often initialized by the weights pretrained on classification datasets (e.g. ImageNet), easily leading to better performance in classification than regression task. Secondly, the existing regression loss functions fully consider the geometric distance of bounding boxes but ignore the influence of regression difficulty. In this paper, we proposed a novel bounding box regression method, which optimizes the regression task from two perspectives. For the feature optimization, we introduce a spatial-attention-based (SAB) layer to the regression head network, making the feature pays more attention to spatial location. For the optimization of loss function, we present a concise method for calculating the penalty term of aspect ratio and construct a new loss based on regression difficulty, which dynamically balances loss scales with different regression difficulties. We incorporate the proposed strategies into RetinaNet and FCOS. Resnet-50 with the pre-trained weight is used as the backbone for training these object detectors. On Pascal VOC, the mean Average Precision (mAP) of RetinaNet is boosted from $81.27 \%$ to $83.12 \%$. The mAP of FCOS is increased from $76.50 \%$ to $78.49 \%$. On MSCOCO, the Average Precision (AP) of RetinaNet is boosted from $33.5 \%$ to $33.9 \%$. The AP of FCOS is improved from $34.9 \%$ to $35.4 \%$. Experimental results manifest effectiveness and generality of the proposed method.

Read the paper · More papers on PaperTik