Decoupling Classification and Localization of CLIP

Muyang Yi, Zhaozhi Xie, Yuwen Yang, Chang Xin Liu, Yue Ding, Hongtao Lu · 2024

Contrastive language-image pre-training, known for its effectiveness in zero-shot retrieval and classification, faces limitations in localization ability, hindering its broader application in diverse vision tasks such as object detection and segmentation. Classifying and locating objects involves gathering global features and local features, respectively. A conflict exists between these abilities when the image encoder updates the weights in pre-training. To address this issue, we propose CLIP-WIRE, which decouples the global and local representation learning by incor-porating registers into the image encoder. CLIP-WIRE allows the model to store global features in the register tokens and focus on local features through patch tokens, thereby jointly improving the classification and localization abilities. Furthermore, a segmentor is designed to validate the enhancement in semantic segmentation tasks. Comprehensive experiments and visualizations prove the significance of the proposed approach.

Read the paper · More papers on PaperTik