A Highly Efficient Zero-Shot Cross-Modal Hashing Method Based on CLIP
Liuyang Cao, Hanguang Xiao, Wangwang Song, Huanqi Li · 2024
With the rapid proliferation of multimedia data, the ability to quickly retrieve relevant information from different modalities of media data is crucial. Cross-modal hashing methods have proven effective in addressing this need. To handle the emergence of new categories, zero-shot cross-modal hashing methods have been proposed. However, existing zero-shot cross-modal hashing methods face challenges in handling multi-label images and require retraining the model as the demand for hash code length changes. In response, we propose a Highly Efficient Zero-Shot Cross-Modal Hashing method based on CLIP (CMZSH-CLIP). Leveraging pre-training on massive paired multimodal data, CLIP's image encoder and text encoder can effectively extract discriminative feature information from different modalities. Based on this, we design cues suitable for multi-label images to guide the model in recognizing the multi-label characteristics of images. Additionally, we introduce a multi-level hash loss function. Under the guidance of this loss function, the hash codes we train can push more discriminative similarity information towards the front of the hash codes. Moreover, our model can be trained once and then freely adjust the desired hash code length as needed. Extensive experiments on three public datasets demonstrate the superiority of our approach.