Active Learning for Open-Set Object Detection with Large Multimodal Model

Zhipeng Zhang, Xiaohang Yuan, Wenting Ma, Jinman Lin, Lei Yang, Meng Guo, Xia Zhao · 2024

Recently, open-set detection based on vision-language pre-training large models has garnered significant attention. Despite significant progress in solving the open-set detection problem, it has been observed that the few-shot or zero-shot capabilities of these models are not always sufficient. This is attributed to: 1. Lack of large-scale visual language datasets for specific areas such as healthcare, transportation, and industrial defects. 2. The absence of visual and language encoders capable of handling irregularly shaped objects in complex environments. 3. The imbalance between positive and negative samples. We explore active learning for open-set object detection with large vision-language pre-training models. While text information effectively conveys the abstract concept of common objects, visual information excels in illustrating novel objects through concrete visual representations. Recognizing the complementary strengths and weaknesses of both text and visual information, we combine text and visual uncertainty to identify the most informative samples for labeling, which can significantly enhance model performance. Our experimental results demonstrate the validity of the method.

Read the paper · More papers on PaperTik