EdgeCLIP: Injecting Edge-Awareness Into Visual-Language Models for Zero-Shot Semantic Segmentation

Jiaxiang Fang, Shiqiang Ma, Guihua Duan, Fei Guo, Shengfeng He · IEEE Transactions on Circuits and Systems for Video Technology · 2025

Effective segmentation of unseen categories in zero-shot semantic segmentation is hindered by models’ limited ability to interpret edges in unfamiliar contexts. In this paper, we propose EdgeCLIP, which addresses this by integrating CLIP with explicit edge-awareness. Based on the premise that edge variation patterns are similar across both seen and unseen class objects, EdgeCLIP introduces the Contextual Edge Sensing module. This module accurately discerns and utilizes edge information, which is crucial in complex border areas where conventional models struggle. Further, our Text-Guided Dense Feature Matching strategy precisely aligns text encodings with corresponding visual edge features, effectively distinguishing them from background edges. This strategy not only optimizes the training of CLIP’s image and text encoders but also leverages the intrinsic completeness of objects, enhancing the model’s ability to generalize and accurately segment objects in unseen classes. EdgeCLIP significantly outperforms the current state-of-the-art method, achieving a deep impressive margin of 17.5% on COCO-20i datasets. Our code is available at github.com/aqingaqinghh/EdgeCLIP.

Read the paper · More papers on PaperTik