CT-FSOD: CLIP-Based Text-Guided Few-Shot Object Detection
Hyugjae Chang, Woojin Lee, Munchurl Kim · IEEE Access · 2025
Few-Shot Object Detection (FSOD) has recently gained significant attention due to its potential to detect objects of novel classes with only a few annotated examples. Existing FSOD methods rely largely on visual features extracted from limited support images of novel classes, which are highly sensitive to background noise, pose variations, and intra-class diversity, thereby limiting generalization performance on object detection for the novel classes. To address these challenges, we propose a Contrastive Language-Image Pretraining (CLIP)-based text-guided FSOD framework, denoted as CT-FSOD, that leverages a CLIP-based vision-language pretraining model to enhance class prototypes (class representative vectors) learning and representation alignment between support and query images. Our framework consists of two core branches: (i) Prototype Distillation Branch that extracts semantically aligned features from support images using a frozen CLIP image encoder, producing robust and semantics-aware prototypes; (ii) Text Matching Branch that generates initial prototypes for novel classes using CLIP text features constructed through CoOp-style learnable context tokens, providing pose-invariant and semantically consistent anchors. During training, query image features from a trainable CNN backbone are doubly aligned with CLIP-derived rich and semantic (i) image features via the Prototype Distillation Module and (ii) text features via the CLIP Alignment Loss. These alignments encourage meaningful and generalizable discriminative feature learning, thus significantly improving object detection performance for novel classes under few-shot conditions. Extensive experiments on PASCAL VOC and MS COCO benchmarks demonstrate that our CT-FSOD achieves superior results compared to existing state-of-the-art methods for FSOD.