Few-shot Military Target Detection with Text-to-image and Vision-language Models
Yan Zhang, Jinjin Zhu, Keyi Cao · 2024
Few-shot detection poses significant challenges for traditional object detection methods due to limited labeled data, overfitting, and class imbalance. These challenges are particularly pronounced in the context of military target detection, which involves unique characteristics and scarce labeled samples. To address these issues, we present a two-stage framework for few-shot military target detection using advanced large-scale language models. In the detection stage, we employ the low-rank adaptation of the large language method to fine-tune the stable diffusion text-to-image model, generating high-quality target samples to enhance the completeness of the target dataset. In the recognition stage, we propose a CLIP model fine-tuning method based on prior probability constraints. This method incorporates the probability of target models existing in the background during the fine-tuning process of the vision-language model, thereby improving the accuracy of target recognition. Furthermore, we introduce a multi-knowledge model fusion method, combining the CLIP model and YOLOv8 model for joint target model recognition, further enhancing the accuracy of target recognition. Experimental results demonstrate the effectiveness and robustness of our framework, showcasing its potential for enhancing few-shot military target detection in challenging scenarios.