Few-Shot Image Classification Guided by Large Language Models
Pinzhe Chen, Hao Chen, Yanxi Liu, Longhui Han · Applied and Computational Engineering · 2025
Few-shot image classification remains a challenging task due to the scarcity of labeled data for novel categories. Prompt-based tuning methods, such as CoOp and CoCoOp, have shown promise by adapting pre-trained vision-language models like CLIP to downstream tasks. However, these methods either suffer from limited generalization or lack semantic grounding. In this paper, we propose LLM-PromptTuning, a novel framework that integrates class-level semantic knowledge from large language models (LLMs) into conditional prompt learning. We generate rich, descriptive prompts for each class using an LLM and combine them with image-conditioned prompt vectors produced by a Transformer-based meta-network. These hybrid prompts are fed into the CLIP text encoder, improving the alignment between textual and visual representations. Experiments on Caltech101, Oxford Flowers102, and Food101 demonstrate that our method outperforms CoOp and CoCoOp, particularly in novel class recognition, achieving a new state-of-the-art in prompt-based few-shot learning.