UDG-Prom: A unified dense-guided semantic prompting for cross-domain few-shot image segmentation
Jiaqi Yang, Xiangjian He, Xin Chen, Yaning Zhang, Jingxi Hu, Linlin Shen, Guoping Qiu · Knowledge-Based Systems · 2025
• MAF preserves low-level feature representations, while fusing global and local information to generate robust class-agnostic features. • TA2MP, as a unified feature transformation mechanism equipped with an automatic learnable prompt branch, reduces human reliance and disentangles domain- and class-specific information through contrastive learning. • UDG-Prom integrates the MAF and TA2MP modules to address the CD-FSS task with SAM. • Our model achieves competitive or superior performance compared to state-of-the-art methods on four CD-FSS benchmarks, and its strong generalization ability is comprehensively validated through evaluations on more difficult cross-domain datasets including CT-Lung (medical) and SUIM (underwater). Large Vision Models (LVMs), exemplified by SAM, contain powerful general knowledge from extensive pre-training, yet they often underperform in highly specialized domains. Building large models tailored for each domain is usually impractical due to the substantial cost of data collection and training. Therefore, a key challenge is how to tap into SAM’s strong knowledge base and transfer it effectively to new, domain-specific tasks, especially under Cross-Domain or Few-Shot constraints. Previous efforts have leveraged prior knowledge from foundation models for transfer learning; however, they typically target specific tasks and exhibit limited robustness in broader applications. To tackle this issue, we propose a Unified Dense-Guided Semantic Prompting framework (UDG-Prom), a new paradigm for Cross-Domain Few-Shot Segmentation (CD-FSS). First, a Multi-level Adaptation Framework (MAF) is used for integrated feature extraction as prior knowledge. Then, we incorporate a Task-Adaptive Auto Meta Prompt (TA 2 MP) module to enable the extraction of class-domain-agnostic features and generate high-quality, learnable visual prompts. By combining learnable prompts with a structured model and prototype disentanglement, this method retains SAM’s prior knowledge and effectively adapts to CD-FSS through category and domain cues. Extensive experiments on four benchmarks show that our model not only surpasses state-of-the-art CD-FSS approaches but also achieves a remarkable improvement in average accuracy.