Deep Learning-Based Multimodal Image Retrieval Combining Image and Text

Md Imran Sarker, Mariofanna G. Milanova · 2022

Multimodal learning is omnipresent in our lives. Human absorbs features in different ways, whether through pictures or text. Combining these features in computational science, especially in Image retrieval problems, poses two significant challenges: how and when to fuse them. Most image retrieval systems use images or text data associated with the image. In this paper, we study the image retrieval task, where the input query is an image plus text sentence that describes the image. The system starts a query triggered by input image and text while taking the help of the Transformer model, which puts attention on both modalities and combines embedded features through the feature fusion technique. We proposed a feature fusion layer using modified Text Image Residual Gating in our work. We have used two methods based on the features extracted from the fusion layer. First, we trained K Nearest Neighbor (KNN) algorithm on the training data, and later we used test data to find a similar image. Second, we used the clustering technique and a support vector machine to compute the nearest neighbor points and cluster the center to see a similar image. We found that SVM (Support vector Machine) is more effective from the results, giving an overall accuracy of 92%.

Read the paper · More papers on PaperTik