Text-to-image Retrieval Based on Zero-shot Transfer Learning with CLIP Model and Vector Database

Junfeng Xie, Junying Chen · 2024

With the continuous increase in the number of internet users, the explosion of multimodal data on the web has led to a growing demand for image retrieval. Current text-to-image retrieval systems often rely on keyword matching, which frequently fails to accurately capture users' retrieval needs. This work proposes a text-to-image retrieval method based on zero-shot transfer learning with the CLIP model and vector database, designed to be applied in text-to-image retrieval systems. By leveraging the semantic alignment capability of the CLIP model for text and images, this study effectively addresses the semantic gap problem in text-to-image retrieval. Additionally, the method utilizes the Milvus database for efficient storage and retrieval of image feature vectors, thereby enhancing system stability. Experimental results on the Flickr30k dataset demonstrate that this method improves the accuracy of text-to-image retrieval, verifying its effectiveness. Furthermore, this work extends the proposed method by incorporating the iFLYTEK Spark Cognitive Model to support image retrieval using Chinese text descriptions, further broadening the system's applicability and reliability.

Read the paper · More papers on PaperTik