GCCA: Utilizing Gradient Matching for Open-set Domain Adaptation of Vision-language Models
Mingxi Jiang, Shengyong Xu · Applied and Computational Engineering · 2025
Unsupervised Domain Adaptation (UDA) traditionally focuses on learning domain-invariant representations by aligning source and target domains, yet struggles in open-set scenarios where target domains contain unseen categories. Recent advances in vision-language models like CLIP, which leverage multimodal pre-training on large-scale image-text pairs, exhibit remarkable zero-shot generalization but still face performance degradation under significant domain shifts. In this work, we propose a gradient matching strategy to enhance CLIP's adaptability in open-set UDA. Unlike conventional prompt-based adaptation methods that independently optimize domain-specific parameters, our approach formulates domain alignment as a gradient consensus problem. Specifically, we enforce gradient direction consistency between source and target objectives to reconcile their optimization paths, while applying gradient norm penalization to prevent overfitting to domain-specific noise. This dual mechanism allows CLIP to preserve its inherent cross-modal discriminability while acquiring domain-robust features. Experiments across diverse open-set adaptation benchmarks demonstrate that our gradient matching framework substantially outperforms existing CLIP adaptation baselines, achieving state-of-the-art performance without compromising model stability.