VLP Based Open-set Object Detection with Improved RT-DETR

Guocheng An, Qiqiang Huang, Gang Xiong, Yanwei Zhang · 2024

Despite the remarkable accuracy of traditional object detectors, they are unable to detect novel categories. This paper proposes a method for open-set object detection based on generating pseudo-labels using the Vision-Language Pre-trained (VLP) model. This approach enables traditional object detectors to perform open-set object detection and can be generalized to all object detectors. Additionally, this paper introduces two improvements to RT-DETR. First, replacing the RepC3 in the fusion module with Manhattan Self-Attention (MaSA) to better construct global features. Second, using MPDIoU loss instead of GIoU loss. The results demonstrate that the improved RT-DETR achieves increases of 3.1%, 3.8%, and 3.1% mAP for all classes, base classes, and novel classes on the Pascal VOC07+12 dataset, respectively. Furthermore, the proposed method shows a 1.3% improvement in mAP for open-set object detection (64.6% mAP for novel classes) compared to ZSD methods.

Read the paper · More papers on PaperTik