DE-CLIP: Unsupervised Dense Counting Method Based on Multimodal Deep Sharing Prompts and Cross-Modal Alignment Ranking
Xuebin Zi, Chunlei Wu · Electronics · 2025
With the rapid development of multimodal prompt learning in unsupervised domains, prompt tuning has demonstrated significant potential for dense counting tasks. However, existing supervised methods heavily rely on annotated data, limiting their generalization capabilities. Additionally, unimodal prompt designs fail to fully leverage the complementary advantages of multimodal data, compromising the accuracy and robustness of counting systems. To address these challenges, we propose DE-CLIP, an unsupervised dense counting method based on multimodal deep shared prompts and cross-modal alignment ranking. DE-CLIP constructs hierarchically ordered textual prompts and optimizes the image encoder via cross-modal alignment ranking loss, which enforces rank-aware embedding learning by aligning visual patches with incrementally scaled textual descriptions, thereby enhancing the model’s numerical perception. The text encoder recursively injects visual information across transformer layers, achieving the progressive fusion of textual and visual prompts to improve multimodal representation. Simultaneously, the image encoder interacts deeply with textual prompts at each transformer layer, strengthening the synergy between visual features and textual semantics. A multimodal collaborative fusion module further enables bidirectional interaction between modalities via self-attention and cross-attention mechanisms, enhancing the model’s capability to comprehend and process complex scenes. The experimental results demonstrate that DE-CLIP significantly outperforms the existing supervised and unsupervised methods across multiple dense counting benchmarks, achieving superior recognition accuracy and generalization ability. This validates its exceptional performance and broad applicability in unsupervised settings.