TCFF-Adapter: Text-Driven Adaption of CLIP for Few-Shot Image Classification

Guanlin Du, Hanzi Wang, Xintao Xu, Yan Yan, Xuelong Li · IEEE Transactions on Circuits and Systems for Video Technology · 2025

In recent years, few-shot image classification has achieved substantial progress. Although existing methods have achieved promising performance, the limited availability of training data often leads to the problem of model overfitting. Model overfitting affects generalization and restricts the effective transfer of knowledge to unseen classes. Moreover, existing methods maintain independence between the image and text modalities during the encoding process, lacking mutual collaboration. This limitation restricts their ability to fully exploit task-specific semantic relationships between visual concepts and textual descriptions. To address this challenge, we propose a text-driven cross-modal feature fusion adapter (TCFF-Adapter) for few-shot image classification. TCFF-Adapter introduces two core components: a cross-modal feature fusion module that constructs joint representations by aligning image and text semantics, and a text-driven adapter that optimizes fused features and dynamically adjusts feature weights in a meta-learning paradigm. By integrating multimodal knowledge with parameter-efficient tuning, our method achieves robust generalization to unseen data without requiring additional fine-tuning. Extensive experiments on eight benchmark datasets demonstrate that the proposed TCFF-Adapter significantly outperforms various state-of-the-art few-shot image classification methods.

Read the paper · More papers on PaperTik