Semantic Prototyping With CLIP for Few-Shot Object Detection in Remote Sensing Images

Tianying Liu, Shuigeng Zhou, Wengen Li, Yichao Zhang, Jihong Guan · IEEE Transactions on Geoscience and Remote Sensing · 2025

Few-shot object detection (FSOD) has been proposed to solve the problem of insufficient data for training, and it has drawn the attention of the remote sensing community in recent years. A mainstream type of FSOD method is to generate class prototypes based on the limited samples to help the construction of classification decision boundaries. However, these constructed prototypes may be far away from the true class centroids in the few-shot scenario. Recently, the vision-language model (VLM) has shown its powerful ability to align the visual features and text features, which leads to strong zero-shot performance on various downstream computer vision tasks when given only texts. Therefore, in this work, we propose to build class prototypes from text descriptions instead of limited visual instances by leveraging a classical pretrained VLM named CLIP. Concretely, we generate prototypes by feeding the CLIP text encoder with class names and enforcing each positive proposal feature to be close to the corresponding prototype. To accelerate the alignment process, we utilize the CLIP visual encoder as another teacher to achieve visual knowledge distillation. Moreover, we adopt prompt tuning to adapt CLIP to the remote sensing scenario. Extensive experiments on two public FSOD datasets, i.e., DIOR and NWPU VHR-10.v2, demonstrate the effectiveness of our method, which yields competitive results with that of existing approaches.

Read the paper · More papers on PaperTik