BOAformer: Object Detection Network Based on Background-Object Attention Transformer
Yong Yang, Chenbin Liang, Shuying Huang, Xiaozheng Wang · 2024
At present, deep learning has achieved great success in the field of object detection. To ensure that positive samples in the image are not missed, most deep-learning object detection methods set many prediction boxes and use the same prediction operation on them. Although these methods can obtain higher predictive values for high certainty objects, they may identify objects with lower certainty, such as small or fuzzy objects, as negative samples or miss them. To address this issue, this paper presents an object detection network based on background-object attention transformer (BOAformer), which improves the accuracy of object detection by establishing the relationship between background and object features. Firstly, the existing classification model is employed to obtain scores of object prediction boxes. On this basis, a scoring mechanism is designed to filter out high scoring object prediction boxes to reduce the computational complexity of the later network. Then, to improve the accuracy of object detection, BOAformer is constructed, which obtains auxiliary classification values of objects by learning the relationship between background and object features, and adds them to the initial detection results. Finally, many experiments are conducted to verify the effectiveness of BOAformer. The experimental results show that the designed network improves the accuracy of object detection and can detect small or lost objects in the background.