A Language-Prompted Model for Semi-Supervised Object Detection
Ying Zeng, Jialong Zhu · 2024
Object detection is a cornerstone of computer vision and artificial intelligence. Despite the advancements made by previous methods, they typically focus on visual signals, often overlooking the potential benefits of integrating language descriptions. Large-scale pre-trained language models offer rich semantic understanding that can complement visual cues, yet their application in object detection remains underexplored. This paper proposes a novel approach: a language-guided model for semi-supervised object detection. Unlike conventional methods that rely solely on visual features, our approach integrates a language model to encode descriptive information corresponding to images. For annotated data, language descriptions are derived from ground truth, capturing object categories and precise spatial coordinates in coherent sentences. In addition to leveraging labeled data, our method utilizes pseudo-labels for language descriptions on unlabeled data, enhancing training efficiency and model generalization. Experiments on MS-COCO dataset validate the superiority of our approach, which outperforms competitive methods with fewer labels.