Modeling Cross-Modal Semantic Transformations From Coarse to Fine in CLIP

Ziqi Peng, Zhenyu Qi, Yang Hua Cao, Yu Kang, Wenjun Lv · IEEE Transactions on Circuits and Systems for Video Technology · 2025

Vision-Language Models (VLMs) like CLIP have advanced image representation through open-vocabulary semantic alignment. Yet, existing few-shot transfer learning methods largely overlook the intrinsic interdependencies between text and image embeddings, limiting their ability to fully transfer CLIP’s pretrained capabilities. To address this gap, we propose Hyperspherical Interpolation Variational Encoding (HIVE), a novel method for few-shot image classification. Our core idea is to shift away from directly training feature extraction capabilities for downstream tasks, and instead focus on exploring the semantic transformation relationships between upstream and downstream tasks. By modeling semantics from coarse to fine granularity, HIVE enables the transfer of original feature extraction and modality alignment capabilities to downstream tasks. Extensive experiments on eight established benchmarks, including CUB and EuroSAT, validate HIVE’s efficacy, achieving up to 46.2% and 80.0% improvements over the original CLIP in 1-shot and 16-shot classification tasks, respectively. Our work underscores the importance of preserving pretrained geometric constraints while exploiting semantic hierarchies for effective few-shot adaptation, providing a principled approach for vision-language model customization.

Read the paper · More papers on PaperTik