CLIP-Based Image Retrieval: A Comparative Study Using CEITM Evaluation
Annapurna P Patil, Abhinav Benagi, Charith Rage, Dhanyatha Narayan, A. Susmitha, Pragya Paramita Sahu · 2024
The demand for comprehensive semantic under-standing for image retrieval systems is increasing rapidly. Traditional textual searches fall short of harnessing rich visual information, for which advanced image retrieval solutions can bridge the semantic gap. In this study, we finetuned CLIP on the flickr8k dataset and introduced a novel Cosine Enhanced Image Text Matching Framework (CEITM) to evaluate image retrieval tasks. The absence of a testing dataset poses a significant challenge to evaluate the image-text retrieval. The CEITM framework provides a measure for effective retrieval by utilizing the semantic meaning between query and captions in large datasets to derive ground truth and overcome need for impractical manual annotations. This approach provides a method for precision based evaluation. The evaluation starts with extracting and encoding the images and text using a vision language model, computing cosine similarity, establishing ground truth and retrieved truth labels finally to calculate precision and recall. The reliance on cosine similarity enhances the ease of evaluating image-text retrieval models by providing valuable insights into their semantic coherence and robustness. This study also shows the behavior of the Pretrained CLIP model's image retrieval task on Descriptive and Key-based captions from the Flickr8k dataset.