Textual Embeddings are Good Class-Aware Visual Prompts for Adapting Vision-Language Models

Fusheng Hao, Liu Liu, Fuxiang Wu, Qieshi Zhang, Jun Sheng Cheng · IEEE Signal Processing Letters · 2025

Due to the parallel nature of the textual and visual encoders, very little attention has been paid to developing prompt learning by using well-pretrained encoders in a serial manner, in which the low-biased high-level semantic information accessible to each other for these encoders is ignored. In this letter, we find that textual embeddings are good class-aware visual prompts for adapting vision-language models, which leads to a new framework called TVPrompt (Textual embeddings as class-aware Visual Prompts). To eliminate the modal gap between text and vision, we design a bridging module, which integrates textual embeddings and class token to produce class-aware visual prompts. To ensure that such prompts could effectively collect class-relevant information, we further propose using masked attention to block the unnecessary interactions. Experimental evidence on benchmark datasets demonstrates that our TVPrompt achieves competitive efficiency and performance.

Read the paper · More papers on PaperTik